dump_stack_trace() and dump_core_file() both run a script on our own pid
and copy its output to stderr a block at a time. The copy ended on any
write() which did not place the whole block:
if (write(2, buf, ret) != ret) {
// *sigh*
break;
}
write() is entitled to do that. stderr is a pipe when SITL runs under
autotest, so a full pipe shortens a write, and a timer signal can cut one
short with EINTR. Either way the copy stopped silently, part-way
through, with no "end dumpstack.sh output" line to show that anything was
missing.
Seen in CI: a backtrace ended at frame #8, in the middle of the frame
which would have named the caller - the one thing the dump exists to
provide. A truncated core dump goes the same way and is even easier to
miss.
Keep writing until the block is out, retry EINTR on both the read and the
write, and say so if a write really does fail.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
calc_alt_delta_m() returned 0 on failure indistinguishably from a
genuine zero delta. This poisoned the approach_start latch if the
query failed at latch time (silently disabling the descent ramp for
the rest of the approach), and folded a single failed query during
the per-cycle clamp into a target of rtl_alt_delta instead of simply
skipping the clamp for that cycle.
calc_alt_delta_m() now reports success/failure via a bool return and
an out-parameter. The latch only commits on success (valid stays
false to retry otherwise), and the clamp is skipped outright on a
failed cycle. Also stop the terrain-relative branch from falling
through to the absolute-altitude fallback when the terrain query
itself fails, which would otherwise mix an AMSL delta into a
terrain-relative ramp.
Adds QRTLGradualAltDescent: flies out ~900m, triggers QRTL, and
asserts the vehicle does not shed more than 15m of altitude in the
first 15 seconds, then completes a normal VTOL landing.
Adds QRTLGradualAltDescentTerrain, covering the altitude-frame-mixing
bug found during review: flies over real sloped terrain near CMAC and
checks the approach ramp's target altitude starts near the aircraft's
current position instead of jumping by the terrain-height gap between
the approach start point and home. Does not cover the POSITION1/2 fix
from the same PR - reliably triggering that path needs an abrupt,
timed fault injection that would need dedicated harness support to
test safely.
The two message-polling loops in these tests are written as
MessageHook subclasses, matching this file's existing
DetectThrottleSpike convention, rather than hand-rolled recv_match
loops. Also fixes a failing location-ratchet check by replacing
home_position_as_mav_location()/home_relative_loc_ne() with the
frame-aware home_position_as_location()/offset_location_ne().
Verified locally: both tests pass with a real quadplane SITL build.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Previously, once QRTL's approach phase started, the target altitude
snapped straight to RTL_ALTITUDE regardless of how much higher the
vehicle actually was, causing an immediate steep descent while still
far from home. The target now ramps down gradually from the vehicle's
actual altitude at the start of the approach to RTL_ALTITUDE, reaching
it at the same distance-from-home threshold where the existing
RTL_ALTITUDE -> Q_RTL_ALT approach ramp already begins, so the two
ramps join continuously. If the vehicle is at or below RTL_ALTITUDE at
that point, behaviour is unchanged.
The altitude and distance at the start of the approach are latched in
ModeQRTL rather than taken from the fixed wing waypoint altitude
offset, which is cleared as the destination is neared. As for fixed
wing waypoints, ALT_SLOPE_MIN gates the ramp: setting it to zero
disables the gradual descent, and altitude changes smaller than it are
still made immediately; ALT_SLOPE_MIN's description is updated to
mention this.
The approach altitude profile continues through the airbrake,
POSITION1 and POSITION2 stages, the same window in which TECS itself
stays active (QuadPlane::should_disable_TECS() only disables it from
QPOS_LAND_DESCEND on), so the altitude TECS is using is never stepped
by handing back to the generic waypoint altitude target early. Over
sloping ground, the altitude delta driving the ramp is computed in
whichever frame (terrain-relative or absolute) QRTL is actually
following, matching the idiom already used by
Plane::set_offset_altitude_location() for the generic waypoint case.
The far-field portion of the ramp is clamped to never command a climb
back up to the ramp line: it only ever demands descent relative to the
vehicle's actual current altitude. Without this, entering QRTL while
already sinking briskly (e.g. a battery failsafe during a fast
descent) would have the ramp instruct a climb back toward the line
before resuming its own gradual descent. The clamp is floored at
rtl_alt_delta so it hands off to the near-field branch at the same
altitude that branch itself targets, rather than opening a new step at
the dist_rtl_alt_reached boundary.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The terrain handler quietly declined to answer requests for tiles
absent from the cache, leaving the vehicle re-requesting the same
block forever - which surfaced as TerrainMission timing out waiting
for prearm with no hint that the harness was the reason. Make an
unserveable request fail the test with a message naming the tile;
install_terrain_handlers_context(unserveable_requests_fatal=False)
restores the old behaviour, since a real terrain server simply does
not answer for data it does not have.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
TerrainMission intermittently died in CI-style runs with "PreArm:
waiting for terrain data", the vehicle asking for the same block at
-26.59 151.84 with a full mask several thousand times: that is
Kingaroy, wanted because state from LargeMissions' Kingaroy missions
was still aboard, and the terrain handler cannot answer - the tile is
absent from the cache and the harness runs the elevation model
offline. The handler's comment already documents not answering as
the correct behaviour for missing data; the missing data itself was
the problem. Reproduce by running
test.Plane.LargeMissions,TerrainMission with the between-test mission
clear disabled.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
BootInAUTO's mission file carried a waypoint at the ARACE field in
Hungary (47.523048 18.813797) - presumably where the mission was
authored. AP_Terrain prefetches terrain for mission waypoints, so
after BootInAUTO runs the vehicle is left with a pending want for a
tile fifteen thousand kilometres from its CMAC home. That pending
want haunts every later test sharing the boot:
- the next test to install terrain handlers is asked for the tile,
which the tilecache cannot serve, and fails fatally
(LoiterAltQLand, 7 of 11 QuadPlane runs at --parallel=8); and
- pending terrain blocks prearm, so tests downstream on the same
worker fail "Prearm bit never went true" (ShipLanding,
SimBatteryResistance, RudderArmedTakeoffRequiresNeutralThrottle,
PrecisionLanding, RTL_AUTOLAND_1 and others - the sweep's
"QuadPlane p8 cluster").
The whole chain reproduces locally with
test.QuadPlane.BootInAUTO,LoiterAltQLand,PrecisionLanding
and passes with the waypoint moved.
BootInAUTO is location-agnostic - it checks the vehicle climbs to 10m
without wandering - so place the waypoint 420m north-east of CMAC,
inside a tile the suite's tilecache actually holds.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Cover get_time_utc(), set_utc_usec() source priority and rejection
rules, get_utc_usec(), get_system_clock_utc(), get_local_time(),
get_date_and_time_utc() and the date-field conversions.
get_time_utc() cases documenting the intended handling of ignored
values below the largest specified element are #if 0'd out; they fail
before the fix in https://github.com/ArduPilot/ardupilot/pull/33667,
which will remove the guard.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A pexpect timeout raised before the test loop had started - during
init, for example - crashed run_tests' exception handler with an
UnboundLocalError on the loop variable, masking the actual failure.
Bind a placeholder first so such timeouts are reported as what they
are.
This is what turned test.BattCAN's peripheral startup panic into
"cannot access local variable 'test'".
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
test.BattCAN has referenced periph-battmon.parm since it was added in
364452ffc8, but the file was never committed: the battery-monitor
peripheral panicked on startup failing to load its defaults, and the
vehicle timed out waiting for it. Configure the SMBus battery
simulation as the universal peripheral does.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
DriveMaxRCIN drives at full throttle for 30 seconds and then disarms
and returns immediately. Disarming stops the motors but not the
vehicle: the rover coasts on for well over ten seconds.
Under --parallel a worker runs many tests in one SITL session, so the
next test inherits that momentum. SafetySwitch ran next and begins
with "Make sure we don't start moving when safety switch enabled",
asserting wait_groundspeed(0, 0.1, minimum_duration=2) immediately. It
was entering that wait at 1.68m/s and spending its whole 30s budget
waiting for a coast-down it had not caused, failing when it ran out
with 0.6s of the 2s settle accumulated.
Centre the sticks and wait for the rover to come to rest before
finishing. A test is entitled to assume it starts from a standstill.
The 0.2m/s bound is chosen from measurement: coasting from ~15m/s the
approach to zero is asymptotic, and 0.1m/s is not reliably reachable
inside wait_groundspeed's 30s budget - it timed out at 0.14m/s. 0.2m/s
is reached with plenty to spare, and drops SafetySwitch's entry speed
from 1.68m/s to 0.11m/s, which it now satisfies in three samples
instead of twenty-eight.
test scripts / build (astyle-cleanliness) (push) Canceled after 0s
test scripts / build (check_autotest_options) (push) Canceled after 0s
test scripts / build (logger_metadata) (push) Canceled after 0s
test scripts / build (param-file-validation) (push) Canceled after 0s
test scripts / build (param_parse) (push) Canceled after 0s
test scripts / build (python-cleanliness) (push) Canceled after 0s
test scripts / build (shellcheck) (push) Canceled after 0s
test scripts / build (validate_board_list) (push) Canceled after 0s
A locked autotest refused to run but exited zero, so a caller checking
the exit code saw a run which did nothing report success.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
test scripting / test-scripting (push) Canceled after 0s
test scripts / build (astyle-cleanliness) (push) Canceled after 0s
test scripts / build (check_autotest_options) (push) Canceled after 0s
test scripts / build (logger_metadata) (push) Canceled after 0s
test scripts / build (param-file-validation) (push) Canceled after 0s
test scripts / build (param_parse) (push) Canceled after 0s
test scripts / build (python-cleanliness) (push) Canceled after 0s
test scripts / build (shellcheck) (push) Canceled after 0s
test scripts / build (validate_board_list) (push) Canceled after 0s
adds periph_board to the quadplane-can frame so restart_SITL_frame
spawns the AP_Periph companion, driver script install helpers and a
plane test checking that scripting monitors fed from CircuitStatus
match a DroneCAN reference monitor on the same periph battery
sends a CircuitStatus message per active battery backend with
circuit_id of the instance number plus one, allowing testing of the
CircuitStatus lua driver
maps uavcan.equipment.power.CircuitStatus circuits onto scripting
battery monitor instances, giving per-circuit voltage and current
monitoring from CAN power distribution nodes