Made a teardown hang in the SMP suite say where it stopped (#646)

The SMP regression suite times out in CI on threadx_thread_priority_change and
the log carries nothing that says why. The reason the log is empty is
mechanical: test_control_return() opens with fflush(stdout), and that is the
last flush before either the test finishes or it wedges. Everything printed
after it sits in stdout's block buffer, and ctest discards that buffer when it
kills the process at the timeout. So the log ends at the test's own result line
no matter where the process actually stopped.

That is enough to place the hang, if not to explain it. The failing runs print
"SUCCESS!" and then stop, which means every check in the test body ran and
passed, and the wedge is somewhere between that flush and exit(). The run of
30 June shows the same signature before any of this test's waits were bounded,
so the hang is not the unbounded wait removed earlier, and the message added
then for an exhausted cap never appears. Retrying tells us nothing new either:
the suite spends 1000 seconds per attempt, twice, to reproduce the same silent
timeout.

Record how far teardown gets, and bound it. A stage variable is updated at each
step from test_control_return() through test_control_cleanup() to exit(), and a
watchdog thread armed on entry to test_control_return() reports the last stage
reached, the per-core scheduler state, and every thread on the created list,
then exits 99.

The watchdog covers teardown and not the test body, because the test body has
no bounded runtime to hold it to. Several tests here wait on a probabilistic
interrupt window: threadx_thread_wait_abort_and_isr_test has been measured
between 0.34 and 439 seconds while passing. Teardown is a fixed amount of work
that takes milliseconds, so a bound on it cannot turn a slow pass into a
failure. The default is 60 seconds, which also means a wedged run now reports
in one minute rather than burning the 2000 seconds two 1000-second attempts
cost today.

The report is written with write() rather than printf() because a wedged thread
may be holding the stdio lock, and a watchdog that blocked on that lock would
reproduce the silent timeout it exists to replace. For the same reason it reads
the ThreadX globals directly and takes no kernel lock; the values may be torn,
which is acceptable for a post-mortem and cannot deadlock.

One walk in test_control_cleanup() is bounded as well. The loop that steps past
the timer thread and the control thread has no terminating condition of its own
and spins for good if _tx_thread_created_count and the created list ever
disagree, which is one of the shapes the timeout could be taking. It now
reports and stops instead.

Off by default in the sense that matters: stderr stays empty and stdout keeps
its buffering, so output is unchanged on a passing run.
TX_TEST_TEARDOWN_TIMEOUT overrides the bound in seconds and zero disables the
watchdog; TX_TEST_TEARDOWN_TRACE echoes each stage as it is reached and
line-buffers stdout so the surrounding output survives a kill too.

Verified against an injected hang at the point the failing runs stop: the
watchdog fires, names the stage, and exits 99. The dump is already informative,
showing thread 0 left at priority 0 with threshold 0 and inherit 0, the same
priority as the control thread, with core 0's execute pointer still on it. All
five SMP configurations pass 110 of 110 at the parallelism CI uses, in both
quiet and trace modes, and the suite runtime is unchanged.

Only the SMP harness is instrumented. The non-SMP suite has not shown this
hang.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Frédéric Desbiens
2026-08-20 08:51:05 -04:00
committed by GitHub
co-authored by Claude Opus 5
parent a5483f0773
commit 83dfbc3534
File diff suppressed because it is too large Load Diff