100 Commits
Author SHA1 Message Date
Frédéric Desbiens 93387b0a60 Merge pull request #783 from eclipse-threadx/dev
Merged changes ahead of the v6.5.2.202603 release
2026-10-02 10:42:44 -04:00
Frédéric Desbiens 28281cd564 Stopped the RISC-V install pulling a desktop media stack (#789)
The RISC-V job installs qemu-system-misc, ninja-build and cmake under a
two-minute per-attempt budget. With recommends enabled that install fetches 67
packages and 87 MB of archives, 388 MB unpacked, because qemu-system-misc
recommends gstreamer, pulseaudio, v4l and a set of codec libraries. Twenty-six
of the 67 are that media stack, and the job runs qemu-system-riscv64 headless.
On a slow mirror the download does not finish inside the budget, and all three
attempts time out.

The install now passes --no-install-recommends. Raising the budget was the
alternative and is not available: TIMEOUT_LONG is already sized so that three
attempts fit inside the ten-minute install step, as its own comment records.

Observed on the run for #783, where all three attempts timed out while a second
RISC-V run at the same minute succeeded, which is the mirror-speed dependence
this removes.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-10-02 10:26:18 -04:00
Frédéric Desbiens 729959fed0 Made the build step timeout a template input (#791)
The build step is capped at fifteen minutes for every consumer, on the same
shape of assumption the install step carried: that a consumer builds once.
Measured build steps are 0.5 minutes for LevelX, 1.1 for FileX and 6.4 for
USBX, so fifteen is right for them.

GUIX builds five configurations in that step, because its suite is staged
across the three steps, and reaches 11.8 minutes. That is 79 per cent of the
cap, and it grows with every regression test added -- six advisories landed
tests this week.

The cap is now an input defaulting to fifteen, matching install_timeout_minutes
above. No consumer changes unless it asks to.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-30 12:42:59 -04:00
Frédéric Desbiens f09846d75a Made the install step timeout a template input (#790)
The install step is capped at ten minutes for every consumer. That suits a
consumer that only installs packages, which is all of them bar one: measured
install steps run 0.2 to 0.5 minutes, and NetX Duo's secure interoperability
job, which builds OpenSSL and wolfSSL from source, takes 2.1 to 2.8.

GUIX does not fit. Its regression suite is staged across the install, build and
test steps to keep each inside its limit, so the install step also builds a
configuration, and on a slow mirror the packages and the ThreadX checkout leave
too little of the ten minutes for it. Three runs on dev were lost that way.

The cap is now an input defaulting to ten, so a consumer that needs longer asks
for it and every other consumer keeps the timeout it has today.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-30 12:00:52 -04:00
Frédéric Desbiens fb4a8761fc Fixed Erbium interrupt masking and PLIC update races
Polling UART output held machine interrupts off for an entire string, which
could lose ThreadX ticks. PLIC enable-word changes could overwrite an update
made by an interrupt handler.

The UART now polls with interrupts enabled and protects only the final status
check and byte write. PLIC enable and disable changes are protected across
their read-modify-write. The README explains the new simulator tests.

Both simulator tests failed on the original code and passed with these fixes.
The 100-million-cycle demo run reported all eight threads and five thread 0
wakeups. AI disclosure and port consistency checks passed. No silicon run.

Assisted-by: Codex (GPT-6-Sol) <noreply@openai.com>
2026-09-30 10:49:55 -04:00
Frédéric Desbiens f19d8c267d Added simulator tests for Erbium interrupt regressions
Long polling UART writes can mask timer interrupts across several tick periods.
PLIC enable-word updates can overwrite a change made by an interrupt handler.
The board example had no simulator tests for either case.

Added two simulator images and a runner. One compares hardware timer progress with
ThreadX ticks after a long print. The other toggles one PLIC source in timer
context while a thread changes another. Both tests fail on the PR code and pass
with the corresponding local fixes.

GCC 15.2.0 and Ninja built both images. erbium_emu reproduced both failures;
control runs passed both tests. Port consistency checks passed. No silicon run.

Assisted-by: Codex (GPT-6-Sol) <noreply@openai.com>
2026-09-30 10:49:55 -04:00
Frédéric Desbiens 0a52df8f87 Stopped the version pass skipping ports in silence (#786)
The port stage matched a version only where "Version" was followed immediately by
four dotted numbers. A port writing three numbers, or a stray letter before the
first, was passed over and kept its old release, and the pass reported success
either way.

The RISC-V32/IAR port is one of them. It reads "Version G6.5.0.202601", and the
letter in front of the number has kept it out of every pass since, so it has
advertised 6.5.0.202601 across two releases it was not part of. Its string now
reads the current release.

The pattern accepts three or four numbers and tolerates a leading letter. A check
follows it: any port header that names a release other than the one being
prepared is listed and the pass stops, so a port the substitution cannot reach
fails the release instead of shipping a version that misreports itself.

Verified: every ThreadX port header now advertises 6.5.2.202603. Against a FileX
tree, a port the substitution cannot reach makes the pass name the file and exit
non-zero, and correcting that string lets the pass complete.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-28 19:56:37 -04:00
Frédéric Desbiens 80dfbe7797 Added the release and disclosure passes to the blame ignore list (#782)
Two commits on dev are mechanical and repository-wide, and neither is in the
list. The disclosure normalisation is the second half of a pass whose first
half, #740, is already listed, so blame behaves differently either side of it.
The version pass is the same shape as the two release preparations above it.

Both are now listed with their file counts. The const object names change is
deliberately absent: it alters behaviour, and size is not the criterion.

All thirteen revisions in the file resolve, and git accepts it as an ignore
file.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-28 14:50:27 -04:00
Frédéric Desbiens a7ded6a3b9 Updated ThreadX version constants and port strings to 6.5.2.202603 (#781)
ThreadX still declares 6.5.1.202602 with hotfix 'a', which is the previous
release rather than the one being cut.

prepare_release.sh updated the five version constants in tx_api.h and the
version string in 211 port headers across ports, ports_smp, ports_arch and
ports_module. The hotfix letter is cleared, since the target has none. Nothing
else changed: every line in the port commit carries a version, and no file
outside an inc/tx_port.h was touched.

The port consistency checks passed before the branch was cut, which is what the
script gates on. Host regression 113/113 with zero warnings, and the AI
disclosure check passes.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-28 14:36:10 -04:00
Frédéric DesbiensandTilen Majerle cf577c7137 Made ThreadX object names const-qualifiable behind an option (#761)
Fixes #61

Object names are exposed as writable pointers throughout the kernel API, which
rejects string literals in C++ and lets a caller modify a name an object still
holds. Information services return those names through writable double
pointers.

Create services, control blocks, information services, the module manager and
trace registration now preserve const qualification, behind
TX_ENABLE_CONST_NAMES. The option defaults to off, so a build that says nothing
gets exactly the types it got before. It is opt-in rather than opt-out because
it changes the type of a public struct field: application code that copies a
name into a writable CHAR * stops compiling, which is a reasonable thing to ask
of a minor release and not of a patch one. Issue #780 tracks making it the
default in 6.6.

Two things the option reaches that its own call sites do not.
TX_CHAR_TO_UCHAR_POINTER_CONVERT has exactly two users, both of them reading an
object name in _tx_trace_object_register, and every form of that macro but the
MISRA one casts the qualifier away without saying so; the conversion is now
const in and const out, so nothing launders const to make the build pass. The
FreeRTOS adapter holds the name pcTaskGetName retrieves in a TX_NAME_CONST
pointer so that it tracks whichever declaration tx_thread_info_get has, and
keeps its writable return type through an explicit MISRA C:2012 Rule 11.8 cast,
because that signature is part of the FreeRTOS API.

Default build: all seven host configurations and all five SMP configurations
build with zero warnings and pass -- 113/113 on five host configurations,
100/100 on the two MISRA builds, 118/118 on SMP, 3/3 FreeRTOS. With
TX_ENABLE_CONST_NAMES set, the host default and both MISRA configurations, the
SMP trace configuration and the FreeRTOS adapter build with zero warnings and
pass.

Co-authored-by: Tilen Majerle <tilen@majerle.eu>
Assisted-by: Codex (gpt-6-astra) <noreply@openai.com>
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-28 14:08:13 -04:00
Frédéric Desbiens 83361990cb Normalized the AI disclosure comment and added a check that keeps it so (#762)
* Normalized the AI disclosure line across every tracked file type

The repository-wide pass covered source files only, so build files, CMake
toolchain files, shell scripts and the GDB and manifest files kept the older
per-edit form of the disclosure comment, which names a product and a model
version. The Cortex-R52 module manager port then merged after that pass and
brought the old form back into the sources as well.

Replaced it with the fixed text in all of them, using the comment character
each file already uses.

Comment-only. 143 files, one line each. The repository now holds 601 files
carrying exactly one disclosure line, none carrying the old form, and none
carrying more than one.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Added a check that keeps the AI disclosure comment in its one accepted form

Nothing enforced the disclosure convention, so the drift it exists to prevent
returned twice: once when a port merged after the normalisation pass carrying
the older per-edit form, and once because that pass had covered source files
only, leaving build files and scripts untouched for months.

Added scripts/check_ai_disclosure.sh, which rejects the superseded per-edit
form, a doubled comment marker, more than one disclosure line in a file, and
any spelling of the line that is not exact. It runs from repo_checks.yml, a
workflow with no path filter, because a source-path filter is what hid the
build files the first time.

The check passes on this branch. Each of its four rules was confirmed to fail
on a tree with that defect reintroduced, and to pass once it was removed.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Normalised the disclosure lines that landed after this branch

Thirteen files reached dev after this branch was written, each carrying the
superseded per-edit form. txm_module_manager_dispatch.h reached it with six
stacked copies, naming the same product and the same model every time -- the
accumulation the fixed text exists to prevent.

Each of those files now carries one disclosure line in the accepted form. Where
the accepted line was already present, the superseded ones are deleted rather
than converted, so no file gains a second.

check_ai_disclosure.sh reported eighteen hits across thirteen files before the
pass and passes after it.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Exempted Markdown from the near-miss disclosure check

The near-miss rule flags any line carrying the phrase "AI assistance" that is
not the accepted text, which is right for source but wrong for documentation.
The contribution guide has to quote the accepted line and say when it applies,
so the check reports two paragraphs of prose as drift and fails the build.

Markdown is now exempt from that rule alone. The three rules that matter for a
documentation file -- superseded form, doubled comment marker, duplicate line
-- still scan it, so a stale disclosure in a Markdown file is still caught.

The check passes against a tree carrying the rewritten contribution guide, and
still fails when a near miss is planted in a source file.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Corrected the trailer command named in the disclosure check

The header comment points a reader at git's own trailer parser to find out
which agents have touched a file. That parser reads trailers only from a block
at the very end of a message, so a squash merge -- which concatenates a
branch's messages -- buries every trailer but the last one mid-message, and an
indented trailer is skipped outright. On dev it finds 122 attributions where
163 exist.

The comment now names count_assisted_by.sh, which reads whole bodies.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-28 14:03:53 -04:00
Frédéric Desbiens a752784c65 Updated ThreadX contribution guide for current workflows (#776)
* Updated ThreadX contribution guide for current workflows

The existing guide left contributors without current build, test, and
submission guidance.

The guide now covers the shared project process and ThreadX-specific C99,
port, simulator, CI, attribution, and release practices.

Documentation links resolve and git diff --check passed. Build tests were
not run for this documentation change.

Assisted-by: Codex (gpt-6-sol) <noreply@openai.com>

* Added visible spacing between contributor setup steps

The ECA and Git setup items had no visible gap in GitHub Markdown.

An indented line break now separates them while preserving list numbering.

GitHub Markdown rendered the items in one list with the visible gap, and
git diff --check passed.

Assisted-by: Codex (gpt-6-sol) <noreply@openai.com>
2026-09-28 14:03:25 -04:00
Frédéric Desbiens 9bf9978e95 Silenced the SMP MISRA shim's implicit conversions (#751)
common_smp/src/tx_misra.c performs four implicit conversions that its
monoprocessor twin makes explicit, and ignores a parameter the twin casts to
void. Built with the warning set common_smp uses, the file reports five
diagnostics that common/src/tx_misra.c does not:

  tx_misra.c:49  conversion to 'int' from 'UINT' may change the sign of the result
  tx_misra.c:93  conversion to 'ULONG' from 'int' may change the sign of the result
  tx_misra.c:151 the same
  tx_misra.c:221 the same
  tx_misra.c:624 unused parameter 'status'

The two files are otherwise the same shim, so this is drift rather than a
deliberate difference: common/src already carries (INT) on the memset value,
(ULONG) on the three pointer differences, and (VOID)status.

That the shim is the file reporting them is the reason to fix it. It exists to
route the kernel's pointer conversions through functions a MISRA analysis can
account for, and an implicit signed-to-unsigned conversion inside it is the
class of construct it was written to remove.

The test trees hide it. test/tx builds the kernel with -Werror, so the
monoprocessor copy could never have regressed this way; test/smp has -Werror
commented out, so the SMP copy warns and builds.

Applying the monoprocessor form to all five sites. Object code is byte
identical before and after, compared with --strip-debug at -m32 -std=c99 with
TX_MISRA_ENABLE defined, so this changes diagnostics and nothing else.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-28 14:02:49 -04:00
Frédéric Desbiens f65727f585 Added a script that counts Assisted-by attributions completely (#775)
Git reads trailers only from a block at the very end of a commit message. A
squash merge concatenates the branch's messages, so every trailer but the last
ends up mid-message, and an indented trailer is skipped as well. Both forms are
real attributions and both are invisible to the trailer parser, which is what
the project has been using to answer which agents touched a file.

The script reads whole commit bodies instead. It groups by value, or lists one
line per commit with --list, and takes a revision and a path the way git log
does.

On dev it reports 163 attributions where the trailer parser reports 122, so a
quarter of the record was unreachable. The difference is entirely squashed and
indented trailers, not new commits.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-28 12:47:25 -04:00
Frédéric Desbiens c805d9d3df Linked ThreadX readers to published documentation (#774)
The ThreadX README sent readers to archived Markdown documentation.

I linked the overview, Modules, and SMP guides to published HTML pages. The general
documentation entry now follows the latest release.

All replacement URLs returned HTTP 200; git diff --check passed.

Assisted-by: Codex (gpt-6-astra) <noreply@openai.com>
2026-09-28 12:46:55 -04:00
Frédéric Desbiens 6098ce67d9 Added a blame ignore list for the mechanical header and version passes (#763)
Running git blame on a file header returns a tree-wide copyright, disclosure or
version-constant pass rather than the commit that last touched the code. There
have been 11 such passes in this repository since 2024, and every header line
in the tree now points at one of them.

Added .git-blame-ignore-revs listing those commits. GitHub applies the file to
its blame view automatically; locally it takes one git config command, which
the file documents in its own header.

Additive. No source file is touched, and every listed commit was verified to be
an ancestor of dev.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-28 12:46:17 -04:00
Frédéric Desbiens 990faf670d Enabled execution profiling for Cortex-R5 (#766)
The Cortex-R5 assembly guarded its execution-profile hooks with only the legacy
TX_ENABLE_EXECUTION_CHANGE_NOTIFY symbol. The documented
TX_EXECUTION_PROFILE_ENABLE configuration initialized profiling without recording
thread or interrupt transitions.

I made all AC5, AC6, GNU, Green Hills, and IAR hooks accept both symbols. I also
extended the port consistency and GNU/LLVM feature checks to cover the current
configuration.

All 849 base assembly files and all 219 TX_EXECUTION_PROFILE_ENABLE files passed
with GCC 14.2.1 and clang 22.1.0. A CMake/Ninja Cortex-R5 profile build emitted
all seven expected hook relocations. Proprietary toolchains were not run.

Assisted-by: Codex (GPT-5) <noreply@openai.com>
2026-09-28 12:45:09 -04:00
Frédéric Desbiens c3bf5e6db5 Made the SMP pt test's randomised phase set preemption thresholds (#760)
The randomised phase of threadx_smp_random_resume_suspend_exclusion_pt_test
guards its preemption-threshold setup with tx_thread_priority > 40, but the
test creates its 1024 source threads with priority = i reset whenever it
reaches TX_MAX_PRIORITIES, which is 32 on the Linux port the regression suite
runs. The guard is therefore never true. The one thing that distinguishes this
test from its non-pt twin in that phase never executes, so the randomised phase
exercises no preemption threshold at all and the test duplicates its twin
there. The threshold the earlier rebalance section sets is unaffected.

Both the guard and the threshold gap are now derived from TX_MAX_PRIORITIES, so
the phase sets a threshold whatever the port's priority count is. On a
32-priority port that is priority > 16 with a threshold eight levels better,
which keeps the original shape of a threshold set on the worse half of the
priority range.

Verified under gdb that the block is now reached, at priority 26 with a
threshold of 18, and that thread_entry's preemption-threshold invariant is
entered for a randomised-phase thread; neither happened before the change.
40/40 passes, 20 runs each in disable_notify_callbacks_build and
stack_checking_build.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-28 12:44:18 -04:00
Frédéric Desbiens 2886bdff6f Stopped the SMP Linux port leaking a critical section on unprotect (#759)
_tx_thread_smp_unprotect takes the Linux mutex on entry and releases it twice when the
protection structure names this core, but only once when it does not, while the matching
_tx_thread_smp_protect took it once either way. _tx_thread_system_return clears the
protection outright, so an unprotect can find it no longer naming its core and return
having released one level fewer than were taken.

A thread that later reaches _tx_linux_mutex_release_all drains that level. The timer
interrupt thread has none on its tick path unless a thread was preempted on core 0, so a
level leaked there while core 0 is idle is never recovered: the nesting count never
reaches zero, the mutex is never handed back to Linux, and every other thread waits on
it for the life of the process while the tick keeps arriving.

The release of the level the matching protect took now happens whether or not the
protection still names this core. Protection bookkeeping and scheduling are unchanged,
and releasing beyond what a thread holds was already a no-op.

Two hung processes captured untraced, in different tests, show the timer thread owning
the mutex with a nesting count of one while parked in its tick wait, the scheduler
blocked in pthread_mutex_timedlock, and the clock advancing 99 ticks a second across a
ten second sample while nothing else changes. On a healthy run that path is taken 0
times in 26,546 unprotect calls. Forcing it on the timer thread leaks one level before
this change and balances after it.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-28 12:43:42 -04:00
Frédéric Desbiens e3ed79dbab Stopped tx_thread_relinquish discarding the rebalance it requests (#758)
_tx_thread_relinquish walks the ready list at the relinquishing thread's priority for a
thread to hand its core to. When the walk reaches one it cannot place, carrying
preemption-threshold or excluded from that thread's mapped core, it sets the rebalance
flag and breaks. That exit falls into the block concluding no other thread is ready,
which is not what the walk found, and that block sets the finished flag. The rebalance
at the end of the function is guarded on that flag, so it never runs.

Every rebalance requested from inside the walk is discarded, because that exit always
reaches the block first. A ready thread behind the obstacle then never reaches a core
while the others hold them and relinquish to each other. The tick keeps arriving, so it
is a livelock rather than a deadlock.

The early return now applies only when no rebalance was requested, which lets control
reach the rebalance path the other three request sites already use.

threadx_smp_relinquish_rebalance_test fills the cores with threads of one priority
relinquishing in a loop, places a thread excluded from every core behind them and a
runnable thread behind that, and fails if the runnable one has not run within 100 ticks.
It fails 3 of 3 before this change and passes 5 of 5 after. Instrumenting a passing run
of threadx_thread_relinquish_test counted 12 rebalance requests from the walk against 4
performed, all 4 from sites outside it. 109/109 in all eight configurations, twice from
clean.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-28 12:43:18 -04:00
Frédéric Desbiens b567428f2a Stopped the Linux port losing a mutex wake-up to a suspend signal (#754)
This is the ThreadX half of the defect fixed for ThreadX SMP in an earlier pull
request. The Linux port suspends a thread with a signal whose handler calls
sigsuspend and does not return until the thread is resumed, and it takes its
critical section with a bare pthread_mutex_lock through tx_linux_mutex_lock. A
thread can therefore be signalled while it is parked on _tx_linux_mutex.

glibc waits for a contended mutex in a loop that re-arms the futex wait after a
signal, and this handler never returns to it. The next unlock hands its wake-up to
that thread, which will not act on it, and any other thread parked on the mutex is
never woken, leaving the mutex free with waiters on it.

tx_linux_mutex_lock now calls a helper that waits with pthread_mutex_timedlock and
retries, so the wait is re-armed every TX_LINUX_MUTEX_RETRY_NSEC and a lost
wake-up costs one retry period instead of the process. The period is one
millisecond. pthread_mutex_timedlock needs _GNU_SOURCE under -std=c99, which both
of this port's build systems already define.

These are the only two ports affected: a sigsuspend-based suspend handler exists
in the Linux and SMP Linux ports and nowhere else, and those two are also the only
ports taking a critical section with pthread_mutex_lock.

Unlike the SMP port, this one has never been observed to deadlock, which is
consistent with one emulated core and far less suspend and resume traffic. The fix
is by inspection, and verified as breaking nothing: all seven configurations pass
with retries disabled, 105/105 in five and 100/100 in two.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-28 12:42:48 -04:00
Frédéric Desbiens 47a5259002 Stopped the SMP Linux port losing a mutex wake-up to a suspend signal (#753)
The SMP Linux port suspends a thread with a signal whose handler calls sigsuspend
and does not return until the thread is resumed. That signal can arrive while the
thread is parked in pthread_mutex_lock on _tx_linux_mutex; the port knows it can,
because _tx_linux_mutex_obtain sets tx_thread_linux_mutex_access around the lock
call for exactly this case, and nothing anywhere reads that flag.

glibc waits for a contended mutex in a loop that re-arms the futex wait after a
signal, and this handler never returns to it. The next release hands its wake-up
to that thread, which will not act on it, and any other thread parked on the
mutex is never woken, leaving the mutex free with waiters on it. That deadlocks
the process: the only thread that can resume the suspended one is the scheduler,
and the scheduler takes this mutex on every pass.

_tx_linux_mutex_obtain now waits with pthread_mutex_timedlock and retries, so the
wait is re-armed every TX_LINUX_MUTEX_RETRY_NSEC and a lost wake-up costs one
retry period instead of the process. The period is one millisecond, half the
scheduler's own idle period. Nothing else changes.

Four hung processes captured untraced, across three tests, show the same state: a
thread in sigsuspend on top of pthread_mutex_lock, the scheduler blocked in
pthread_mutex_lock, and the mutex reading free. Standalone,
threadx_smp_random_resume_suspend_exclusion_test hung 3 times in 100 runs and 2
in 55 before the change and 0 in 400 after it.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-28 12:42:20 -04:00
Frédéric Desbiens a9bd72de21 Merge commit from fork
* Fixed the Cortex-A7 module data check to validate whole ranges

The Cortex-A7 module port received a module instance, a start address and a
byte size from the common Module Manager, then discarded the instance and the
size and asked the MMU to translate the start address alone, for reading only.
A range was therefore accepted whenever its first byte happened to be readable
by the module, so privileged dispatch code could read or write past the end of
the module's mapping, or use read-only module code as a write destination.

The check now takes the size and an access intent. The module's own data region
is answered from the manager's records, which name it exactly and cost no
translations; everything else is answered by translating every page the range
touches with the requested unprivileged access. Empty ranges, ranges whose last
byte would wrap, and ranges that leave the recorded data region partway through
are all refused, and the walk is bounded by TXM_MODULE_MANAGER_DATA_CHECK_MAX_PAGES
so one module request cannot impose unbounded work on the kernel. Translation is
only believed while the requesting module's own context is loaded.

The outside direction gets its own check rather than negating the inside one:
a range that reaches into the module only partway is inside neither answer, and
negating a whole-range check would have let it pass as a kernel object.

Other ports keep their single data check through backward-compatible fallbacks
in the common header; their preprocessed dispatch output is byte-identical.

Regression coverage runs on the host by standing a simulated page map behind the
one architecture primitive, and reaches 100% line, branch and call coverage of
the new logic. Twenty-four of its expectations fail against the previous
implementation.

Assisted-by: Claude Code (Opus 5)

* Drove the Cortex-A7 range check through a live MMU

The host test for this check stands a simulated page map behind the port's CP15
primitive, so it never executes an address translation. That leaves the part
most easily got wrong unexercised: an encoding naming the wrong operation, or a
PAR fault bit read the wrong way round, passes it without complaint.

This adds a bare-metal image that builds a short-descriptor translation table,
enables the MMU, and drives the real check through the real translations on a
real Cortex-A7 translation regime, plus a script that builds and runs it under
an emulator. Sections are mapped for unprivileged read/write, unprivileged read
only, and privileged only, with an unmapped section behind the read-only one so
that a range can be made to leave its mapping partway through.

That last case is the one the replaced check accepted, and the test asserts
both answers against the same live MMU: the current check rejects the range,
and translating only its first address accepts it.

It also confirms what the simulated map could only assume, that ATS1CUW denies
a write to a read-only mapping while ATS1CUR allows the read. The write intent
is the reason the port asks for two translations rather than one.

The script skips with a notice when the cross toolchain or the emulator is
absent, so a machine without them does not fail the build.

Assisted-by: Claude Code (Opus 5)
2026-09-28 11:31:41 -04:00
Frédéric Desbiens 5f9ebde847 Merge commit from fork
A memory-protected module's queue request reaches the Module Manager's dispatch
layer, which decides how much of the module's buffer the privileged queue copy
may touch and hands that extent to the module port's range check. A queue's
tx_queue_message_size is a count of ULONGs and the copy moves that many words, so
the extent is message_size * sizeof(ULONG) bytes. Front-send passed the word count
itself. Send and receive, either side of it in the same header, have always
multiplied.

So a source buffer ending message_size bytes into memory the module may reach was
accepted, and _txe_queue_front_send then read message_size * sizeof(ULONG) bytes
from it in privileged mode: 3 * message_size bytes past the end of what the module
is allowed to read, which is 48 bytes at ThreadX's default message size limit and
96 at the limit this test tree builds with. The words land in the queue, and a
module that can receive from that queue reads them back, so the over-read is
disclosed rather than merely performed. Where the memory past the region is not
mapped for the requesting module, the privileged read faults instead.

The buffer cannot be aimed: it has to pass the same check to be accepted at all,
so it is anchored to the end of a region the module can already read and the
overrun is the fixed window immediately after it. That bounds what this reaches;
it does not make it the module's business.

The fix is the multiplication the other two services do, in the units the port
check has always expected. It changes nothing for a module loaded without memory
protection, because the check it corrects is inside the dispatcher's memory
protection block, and the tests hold that.

The regression test drives the three queue dispatchers directly, which nothing in
the tree did before, and measures the extent each one validates rather than
sampling either side of it: for a given message size it walks the room left
between the buffer and the end of the module's region and finds the smallest
amount the dispatcher accepts. That number is the extent. It has to be
message_size * sizeof(ULONG) for all three services, in the module's data, in a
shared region registered to it, and -- for the two services that read the buffer
rather than write it -- in the module's read-only image, which is where a module
sending a constant message sends it from. A receive destination in the read-only
image is refused at every extent.

Two things had to be arranged for a dispatcher to be testable on the host at all.
The dispatch table is a header of static functions wrapped one per service in
#ifndef TXM_<SERVICE>_CALL_NOT_USED, and at -O0 a compiler emits them all, so
linking the whole table would mean standing up every kernel entry point it
reaches. Defining the 93 guards the test does not exercise compiles the header
down to the three queue services, which leaves four symbols to stub and exercises
that conditional compilation, which nothing else in the tree does. The extent also
has to reach a check that honours it, so the test takes its module port headers
from a Cortex-M port, whose inline data check is the shape the memory-protected
module ports share. The Cortex-A7 port's data check takes the pointer alone and
translates a single address, so on that port no extent reaches anything and a test
built on it would pass equally with and without this change.

That is also the affected-port claim: the defect is in the portable dispatch layer
and was live on every memory-protected module port that honours the size it is
given, which is all of them except Cortex-A7, where the port's own check discarded
the size before this one's error could matter.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-28 11:07:09 -04:00
Frédéric Desbiens ee4c2fb6c7 Merge commit from fork
A memory-protected module's object lookup arrives at the Module Manager's dispatch
layer, which decides how much of the module's name buffer the privileged comparison
may touch. It checked one byte.

_txm_module_manager_object_name_compare reads a character from each name at the top
of every iteration and tests the remaining search length at the bottom, so a declared
length N permits reads at offsets 0 through N: N+1 bytes, the extra one being the
terminator a length excludes. The extended service receives that length in its extra
parameter array, validates the array correctly, and then never uses the length to
bound the buffer it describes. The deprecated service carries no length at all; its
manager wrapper calls the same implementation with the largest value a UINT can hold,
so it removes the bound rather than lacking one.

The comparison stops at the first differing character, so a walk past the end of the
module's memory looks as though it needs the bytes there to match a name the module
chose. It does not. A create service stores the name pointer a module supplies in the
control block rather than copying the string, so a module can register an object whose
name is the very buffer it then searches for. Both sides of the comparison are then
the same address, every character matches by construction, and the walk continues
until it meets a terminator in memory the module does not own and cannot see. That
walk runs in privileged mode with _tx_thread_preempt_disable raised, and none of the
27 memory fault handlers lowers it again, so a fault on it costs more than the
requesting thread.

The fix validates the range the comparison may reach. The extended dispatcher now
checks the name over name_length + 1 bytes, refusing a length whose range cannot be
expressed, and the extra parameter array is checked first because the length comes out
of it. The deprecated dispatcher refuses a memory-protected module outright, because
no range can be derived from a pointer alone; a module without protection is unchanged,
as it is for every other check in this layer.

_txm_module_manager_object_name_compare is left as it is. Reading N+1 bytes for a
declared length of N is the documented contract, and the defect is that the contract
was never checked.

The regression test drives both lookup dispatchers and measures two extents rather than
sampling either side of a boundary, because the finding is the relationship between
them. It anchors the name buffer to the end of one of the module's regions and walks it
backwards to find the smallest room the dispatcher accepts, which is the extent
validated; and it fills memory with a filler character, plants the only terminator at a
chosen depth, and asks for an object named with exactly the bytes up to it, so that the
deepest depth the lookup can be made to return from is the extent read. Against dev's
copy of the two headers the test exits 1 with 408 failed expectations of 1130: a
31-character name with one byte of room is accepted and read 31 bytes past the region,
and the aliased walk reaches every depth offered. With the fix it exits 0 with 1490
expectations and none failed.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-28 10:58:58 -04:00
Frédéric Desbiens a967d005f6 Merge commit from fork
_txm_module_manager_tx_block_pool_create_dispatch proved that two
ALIGN_TYPE words of the extra-parameter array a module supplies lay
inside the module's data, and then used indices 0 through 3. The two
words it had not proved were read twice each in privileged context:
once by the dispatcher's own buffer check, which takes index 1 as the
pool start and index 2 as the length to range-check it over, and once
when the argument list for _txe_block_pool_create was built. So the
out-of-bounds read was performed by the validation code itself, and
the size the caller's pool start was then validated against was a
word the manager had not proved the module owned.

A module reaches this by calling its kernel dispatcher directly with
an array that ends two words before the boundary of its data or of a
shared region, which the module library's own four-word array never
does. The read window is eight bytes on every module port and there
is nothing in it the module can lengthen.

The extent becomes sizeof(ALIGN_TYPE[4]), which is what the two
siblings of the same shape have always had: queue create validates
four words for four, and byte-pool create three for three. Both
checks on this path are data-only, with no fallback to the portable
code-region check, so the four Cortex-A35 and Cortex-A35 SMP module
port combinations whose data check is the constant (TX_SUCCESS) refuse
these create requests from a memory-protected module before and after
this change alike.

The regression test drives all three create dispatchers and measures
the extent each one validates rather than sampling either side of it:
it anchors the array at the end of each region a module owns and walks
it backwards until the dispatcher accepts, and the smallest room
accepted is the extent. It measures the highest index each dispatcher
uses through what the stubbed service records, so the invariant the
three are checked against is measured on both sides. It also measures
what a module can learn from the answers, since the buffer check is a
comparison against a threshold the module chooses and the dispatcher's
refusal is distinguishable from every status these services return.
Compiled against the previous dispatch header the test reports the
extent as two words and the service reached carrying a word planted
past the end of the region; against this one it reports four.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-28 10:52:12 -04:00
Frédéric Desbiens 35d4137126 Merge commit from fork
A module calls txm_module_object_allocate and chooses object_size. The manager
adds sizeof(TXM_MODULE_ALLOCATED_OBJECT) to it and asked the object pool for the
result without checking the addition, so a size near the top of a ULONG wrapped:
object_size 0xFFFFFFF8 asked for eight bytes, the pool served them, and the
manager then wrote its sixteen-byte header -- owner, list links and size -- into
that eight-byte block and linked it into the module's allocation list. Eight
bytes of the next block's contents or header are gone by the time the call
returns, and the shared object pool that every module allocates from is the thing
that was corrupted.

The addition is now made with the repository's own overflow-checked helper, which
was already used on both load paths and simply never reached this one, and it is
made before the protection mutex is taken so a refused request leaves the pool,
the allocation list and its count exactly as they were. Sizes that survive the
addition need no further bound here: _txe_byte_allocate refuses a request larger
than the pool, so the alignment round-up in _tx_byte_allocate is never reached
with a value that could wrap in its turn.

Regression coverage runs on the host by modelling the byte pool behind the
manager. The model applies the two bounds _txe_byte_allocate applies and hands
out a block of precisely the requested length with a guard band immediately
after it, so an undersized allocation is caught as the out-of-bounds write it is
rather than inferred from arithmetic. Thirty-two of its expectations fail against
the previous implementation, twenty reporting state a refused request must not
have touched and two reporting the eight and twelve bytes written past the end of
the block. It reaches 100% line, branch and call coverage of the changed
function; the lines it leaves uncovered are its own failure reporting, plus the
modelled pool's zero-size refusal, which only an unfixed build reaches.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-28 10:37:13 -04:00
Frédéric Desbiens ff4be30656 Covered the control block ID cleared when module memory is given back (#779)
The control block ID that _txm_module_manager_object_deallocate clears on its way
out is read by the size-based object check, which is the one kept for dispatchers
outside this repository. Nothing exercised that path. The expectations covering
the clearing all went through object authentication, which consults the kernel's
created list and would refuse a recycled address whatever its ID said, so they
would have passed with the clearing removed.

A section drives the size-based check directly. A module plants the value of an
ID into memory it owns, gives the allocation back, and is handed the same address
again; the check accepts the planted ID before the deallocation and refuses it
afterwards, which is the difference the clearing makes.

The authentication test holds 125 expectations and passes. Removing the ID
clearing fails it. The full suite passes 107/107 in default_build_coverage with
no warnings.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-28 10:22:07 -04:00
Frédéric Desbiens 2930618cfe Merge commit from fork
* Refused to give back the memory of a live kernel object

A module allocates the control blocks of its kernel objects from the Module
Manager's object pool and can ask for that memory back by address. The manager
released it whatever was in it, including a control block the kernel was still
using. Deallocation is not deletion: the object stays on the created list for its
type, a thread stays wherever it was on the ready or suspension lists and stays
schedulable, and an active timer stays on the timer list. Nothing on those paths
consults a control block ID, so nothing about the memory having been freed stops
the kernel from using it -- a created list walk reads a name pointer out of it and
follows that pointer, a create or delete of another object of the same type writes
through the created links in it, the scheduler switches the stack pointer to the
word at offset 8 of it and pops a saved processor state, and timer expiration
calls the function pointer in it. Meanwhile the byte pool is free to hand those
same bytes to the next allocation, so what the kernel goes on reading as a control
block becomes whatever the next owner of the memory puts there, through ordinary
create services.

The manager now refuses to release memory that holds an object which is still
created, and returns TX_DELETE_ERROR without touching the allocation, the object
or the memory. Deleting the object first is what makes its memory releasable,
which is the sequence the delete dispatchers already follow and the one module
authors are told to follow.

The question is answered from the kernel's created lists, not from the control
block. A control block ID is not evidence that an object is there: an allocation
that was never created can be carrying the value of an ID, and refusing on that
would strand memory a module is entitled to have back. Deallocation is also given
an address and nothing else, so unlike a typed service it cannot be told which
list to search, and each of the eight lists is searched in turn. The search is not
narrowed by the size of the allocation, because that would be sound only if every
object had been created through a size-checked path, and a module running without
memory protection creates objects through no such path -- a queue at the start of a
thread-sized allocation is a case the tests here cover. Each list is searched in
its own interrupts-disabled window, bounded by the count the kernel keeps beside
it, so the longest window is the length of one type's list and a list whose links
have been damaged cannot make the search run on. The whole search is one pass over
the objects the system has created, paid once per object deallocation.

Storage that is not an object is released exactly as before, which is what keeps
cleanup after a create that failed or was abandoned working, and what makes the
release each delete dispatcher performs after a successful delete go through.

The address the request arrives with is now checked before the manager's private
header in front of it is read, rather than partly after. The size of the
allocation comes from that header, so there is nothing to validate a size against
until the header has been read, and the previous order established only where the
header started: a header that began inside the pool and ended past it had its size
word read from outside the pool. That check has moved out of the dispatcher into
_txm_module_manager_param_check_object_for_deallocation, alongside the other
parameter checks the dispatch table uses and where a test can reach it, and it now
also refuses a size that would carry the end of the allocation past the top of the
address space instead of wrapping it, since a wrapped end compares as though the
allocation were inside the pool.

The 301 expectations in the new test drive all eight object types through allocate,
create, a refused deallocation, delete, and a deallocation that succeeds, and
assert after the refusal that nothing reached the pool, that the allocation is
still on the module's list at the head of it, that the control block still carries
its ID and that the object is still live. They cover storage that was never
created, storage carrying nothing but a plausible ID for each of the eight types,
an object at the start of an oversized allocation, every aligned interior offset
of a live object with that object's own ID planted at it, an application-owned
object outside the pool, another module's allocations both live and raw, releasing
the head of a list of several and releasing the same address twice, a type that has
no created list, the bounds on the search against a list longer than its count and
against a count larger than its list for every type, the boundary addresses at both
ends of the pool, a crafted size, and an object pool that was never created.
Removing the refusal fails 43 of them; removing the size wrap guard fails one.

Line and branch coverage of the three new functions and of the changed
_txm_module_manager_object_deallocate is 100%, with one exception that is test
scaffolding rather than product code: the host shim's stand-in for TX_RESTORE has
an underflow guard, and the test asserts that branch is never taken. All 99 tests
pass in each of the five configurations the tree builds with GCC 14.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Removed the object deallocation bounds check superseded by the search

_txm_module_manager_param_check_object_for_deallocation() bounded the private
header in front of a caller's address before the deallocator read it. The
deallocator no longer reads that header: it finds the allocation by searching
the module's own allocation list, which never dereferences the address, so the
bounds check now guards a read that does not happen.

The function, its prototype, its macro and its one call site in the
txm_module_object_deallocate dispatcher are removed, along with the twelve
expectations that covered it and three declarations left unused by their
removal. The search proves more than the check did: the bounds test established
only that the header lay inside the object pool, while the search establishes
that the address is the exact start of one of this module's allocations.

The live object deallocation test holds 289 expectations and passes. Reverting
the live-object guard still fails 51 of them, the same 51 as before the
deletion, so nothing the removed expectations covered was load-bearing. The
full suite passes 108/108 in default_build_coverage, with no warnings.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-28 10:13:22 -04:00
Frédéric Desbiens 3276b0efd5 Merge commit from fork
A module reaches a kernel object's delete service through the Module Manager's
dispatch layer, which asked one question about the pointer: does the object lie
outside the module's own code and data. That is the policy for using an object,
and using an object a module does not own is supported and intended --
txm_module_object_pointer_get_extended searches the system's created objects by
name and hands back objects the application created and objects other modules
created, so that a module can send to a shared queue, take a shared semaphore or
get a shared mutex. Destroying one of those is not sharing it. A module could
name any object it had a pointer to and have the privileged dispatcher delete it:
a host service lost its queue, waiters were resumed with TX_DELETED, an owned
mutex's priority inheritance was unwound, and an active timer stopped.

Nothing in the manager intended that. The deallocation the delete path performs
afterwards has always required the memory to belong to the calling module, so a
delete of anything else could only ever end in TX_PTR_ERROR -- the ownership
requirement was already there and was enforced one step too late, after the
kernel object had been irreversibly destroyed and its waiters woken. An error
returned at that point describes a cleanup failure and restores nothing.

The eight delete dispatchers now ask the question before the delete rather than
after it. A memory-protected module may delete an object only if the object is
one it allocated from the manager's object pool, at the exact address the manager
returned to it, for an allocation made for that type's control block size -- the
same conditions the create dispatchers already apply, since deletion is the
inverse of creation. Anything else returns TXM_MODULE_INVALID_MEMORY, which is
what every other parameter rejection in the dispatch table returns, so no new
value enters the module ABI and a module cannot use a delete request to tell
"not yours" apart from "not a usable address". The object stays created, nothing
waiting on it is resumed, and no created count moves, because the kernel was
never asked.

The ownership question is deliberately separate from whether an object is there
at all and of the expected type. It establishes nothing about liveness or type,
and it is composed with the checks that do rather than replacing them.

_txm_module_manager_object_deallocate reached the manager's private header by
subtracting from whatever address the caller supplied, and read it before
deciding whether it was a header. All eight delete dispatchers call that function
directly, so an object the module does not own arrived there as a matter of
course rather than exceptionally: for an object the application allocated
statically, or for any address outside the object pool, the words in front of it
were unrelated memory read in privileged mode -- and acted on, since an address
whose preceding words happened to name the calling module was unlinked from that
module's allocation list and handed to the byte pool.

It now finds the allocation by searching the module's own allocation list. The
caller's address is compared and never dereferenced, so an address that names
none of this module's allocations is refused without a privileged read of
anything in front of it, and the search establishes what the header read could
not: that the address is the exact start of an allocation rather than somewhere
inside one. The search is bounded by the count the manager keeps beside the list
and runs with interrupts disabled, so a damaged list cannot make it run on and it
cannot observe the list being changed under it. The fix is in the deallocation
itself rather than at a call site, so it covers all eight delete dispatchers, the
deallocation request a module can make directly, protected and unprotected
modules alike.

txm_module_manager_stop needs no exemption and was not given one. It deletes the
objects a module created by calling the internal _tx_*_delete services directly,
and identifies them with _txm_module_manager_created_object_check rather than
through a module request, so no caller-facing check stands in its way.

The 209 expectations in the new test drive all eight object types through
allocation, an ownership check by the owning module and by another, a check
against a larger and a smaller expected size, and a deallocation that succeeds.
They cover an object the application owns, deliberately preceded by a header
naming the requesting module so that reading it can be seen to have happened;
another module's object, checked and refused from both sides; every aligned
interior offset of an allocation and the addresses of its header, its end, the
pool's ends and one past the pool; a null pointer and an address with no room for
a header in front of it; a request from no module at all; memory that has changed
hands, where a stale pointer names an address that now belongs to another module;
the bounds on the search, against a count smaller than the list and against a
count larger than a list whose links are broken; releasing the head, the middle
and the last of a list of three; and an object pool that was never created. Every
call asserts that the interrupt lock came back balanced, and that the search ran
with interrupts disabled exactly one deep.

Restoring the previous deallocation fails four of them and then segmentation
faults, on the null-pointer case, where the header in front of address zero is
read; removing the ownership check fails 74.

Line and branch coverage of the two new functions and of the changed
_txm_module_manager_object_deallocate is 100%, with two branch outcomes excepted
that are test scaffolding rather than product code: the host shim's stand-ins for
TX_DISABLE and TX_RESTORE carry a maximum-depth test and an underflow guard, and
the real ports' primitives are inline assembly with no branch at all. All 99
tests pass in each of the five configurations the tree builds with GCC 14.

Nothing in the tree compiles the dispatch header, so the eight changed
dispatchers were cross-compiled by hand: the two changed C files and a
translation unit that includes the header build at -Werror with arm-none-eabi-gcc
13.2.1 for every GNU 32-bit module port -- Cortex-A7, M0+, M23, M3, M33, M4 and
M7 -- and produce the same -Wall -Wextra warning counts as before the change, 117
on Cortex-A7 and 116 on each of the others. Cortex-R4 and RXv2 have no GNU module
port, and the two AArch64 module ports have no toolchain available here; the new
arithmetic is a comparison of two addresses of the same type, so a wider
ALIGN_TYPE changes nothing about it.

MISRA: all if bodies are braced, the address comparison goes through ALIGN_TYPE,
no pointer the caller supplied is dereferenced, and no goto appears. No deviation
is required.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-28 09:23:49 -04:00
Frédéric Desbiens 7495b02374 Merge commit from fork
A memory-protected module names the kernel objects it wants operated on by
address, and the Module Manager decided whether a privileged service could
dereference that address by asking only whether it fell outside the module. The
object pool is outside every module, so that test was satisfied by an address
shifted into the interior of one of the module's own privileged allocations,
which denotes no object at all. The bytes such an address presents as a control
block are bytes the module put there through ordinary create and set services,
so the control block ID at the front of them could be made to read as any type
the module chose, and the _txe_ layer's ID test then agreed. A module could
therefore have the kernel read and write fields of an object that does not
exist, at an address it picked; the reported chain reaches a privileged memset
across an attacker-chosen range that way.

Authentication now answers the question the location test could not: is this the
exact address of a live kernel object of the type this service expects. It is
answered from the kernel's created list for the type, which the create and delete
services maintain, so membership establishes at once that the address is an
object start rather than an address inside an object, that the object is of this
type, and that it has not been deleted. That list is also what makes an object
the application created and shared with a module authenticable at all, since the
manager allocated no such object and has no record of its own to consult, so
sharing keeps working and needs no new registration API.

Nothing a module can influence is used to establish a type. An earlier form of
this fix accepted an address that was the exact start of one of the calling
module's own allocations of the right size and then took the type from the
control block ID, which is cheaper -- a module's allocation list is much shorter
than the system's created list for a type. The tests here refused it: a byte
pool, a mutex and a timer are all the same size on a 32-bit target, so an
allocation created as one of them and presented as another passed both the
address and the size test, leaving the ID as the only thing between the module
and a type confusion. The ID is still checked, after the created list has settled
the question, because it is the test the _txe_ services make and it is compiled
away with them under TX_DISABLE_ERROR_CHECKING. An address a module named is not
read at all until a kernel record says there is an object there.

All 58 object-using dispatchers now pass the object type rather than a control
block size, since a size cannot establish a type. Both scans are bounded by the
counts the kernel and the manager maintain beside their lists, so the cost of one
check is bounded and a list whose links have been damaged cannot make a scan run
on, and both run with interrupts disabled, which is the protection those lists
are maintained under and which adds no blocking point or priority inversion to a
kernel request. The cost is a walk of the created list for the type, paid only by
memory-protected modules; the size-based check is kept for dispatchers outside
this repository and hardened to refuse object pool interiors, which is the part
of the attack it can see without a type.

Two further changes close paths the authentication alone would leave open. Thread
reset now refuses a thread that carries no module instance: such a thread is
authentic, can be found by name and satisfies reset's own state test, and reset
reads the shell entry function out of that instance, so a null one was followed
in privileged mode. Object deallocation now clears the control block ID as it
gives memory back, so an object freed without being deleted first stops being
vouched for at the moment the manager stops owning its memory, rather than
returning to the pool still carrying a valid ID for the next allocation placed
there to present.

The 120 expectations in the new test offer every aligned interior offset of a
legitimate object as a foreign type, with that type's own ID planted at the
offset, and assert that none is accepted; they cover all eight object types
against each other, uncreated allocations, deleted objects, freed addresses that
have been handed out again, objects the application owns and memory that merely
carries a plausible ID, objects another module allocated, and the bounds on both
scans. Reverting the object checks to the permissive policy fails 89 of them;
removing the thread reset guard makes the test die with SIGSEGV inside the reset
path; removing the ID clearing in deallocation fails 2. All 99 tests pass in each
of the five configurations the tree builds with GCC 14.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-28 09:08:55 -04:00
Frédéric Desbiens d55fe1ef2f Stopped the release script adding a Co-authored-by trailer (#777)
prepare_release.sh committed the version passes with a Co-authored-by trailer naming an
AI. That trailer asserts authorship an AI cannot hold: the human contributor signs the
ECA and is solely responsible for the contribution. It is already in the published
history, on the 6.5.1.202602 and 6.5.1.202602a preparations.

The commits now carry their subject alone. A version pass is mechanical sed output, so
no agent produces it at run time; an agent that runs the script records its own
Assisted-by trailer on that run instead. The script also gains the AI disclosure line it
was missing.

Ran the patched script against a scratch clone targeting 6.5.2.202603. Both commits come
out with an empty trailer block, the version constants and 211 port version strings
update as before, and check_ports.sh passes.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-28 07:28:23 -04:00
Frédéric Desbiens b37cd4a81a Fixed RISC-V regression portability failures (#773)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
RISC-V regression builds ran only at -O0, leaving a timer callback counter that
spins forever at -O2. The trace regression was excluded because it referenced
a port-specific interrupt-save variable, and its adjacent pools could start
misaligned on RV64.

Made the timer counter volatile, gave the trace test aligned pool storage and a
portable saved-interrupt value, and enabled it on RISC-V. Added an -O2 QEMU
configuration to keep the optimized failure covered.

CMake/Ninja/QEMU: RV64 default, optimized and trace suites passed 97/97 each;
RV32 passed 96/96 each. Two ISR event tests passed 30 repeats each at -O2.
The Linux/GCC 14 trace test passed. The reported timing resonance did not recur.

Assisted-by: Codex (gpt-6-sol) <noreply@openai.com>
2026-09-23 16:52:44 -04:00
Frédéric Desbiens 70a5300977 Added a ThreadX module manager port for the Cortex-R52 (#639)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
r52_fvp / r52 (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
Added the first GNU ThreadX module port for an Arm R-profile core.

  The port combines the PMSAv8-R MPU model from Cortex-M33 with the AArch32
  privilege and processor-mode handling from Cortex-R4. It provides the complete
  module path: headers, manager sources, scheduler integration, user-mode entry,
  SVC dispatch, data- and prefetch-abort capture, fault notification, relocatable
  module loading, and shared-memory regions.

  Module isolation uses eight MPU regions (8–15) with a 64-byte granule. The
  scheduler replaces those regions on each module switch and manages a separate
  privileged loading window in region 16. Assembly-visible structure offsets and
  region-layout assumptions are checked at build time.

  Added independent module demonstrations for the S32Z280-594EVB and Armv8-R AEM
  FVP. The automated FVP regressions cover:

  - Loading the same position-independent module at different addresses
  - User-mode data and instruction access violations
  - Fault capture and notification for both abort types
  - Shared-region access, alignment, exhaustion, empty-size, and overflow handling
  - Required module-property combinations
  - The GCC CLZ-based priority search

  The port requires user mode and memory protection together, at least 17 EL1 MPU
  regions, and currently validates A32 modules; Thumb module execution remains
  unvalidated.

  Validated with GNU Arm 14.3.1. The FVP suites pass 8/8 in the default
  configuration and 11/11 with hard-float, FIQ, and interrupt nesting enabled.
  The S32Z280 images build cleanly, and the module isolation and relocation paths
  were exercised on S32Z280 silicon during development.

  Matching user documentation is provided by rtos-docs-asciidoc PR #41.

  Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-17 15:39:30 -04:00
Frédéric Desbiens ad558a7b1f Fixed the zero trace time stamps in the Linux ports' MISRA builds (#749)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
Both Linux ports define TX_TRACE_TIME_SOURCE as _tx_misra_time_stamp_get() when
TX_MISRA_ENABLE is set, and neither implements that function, so both inherit
the generic `return(0);` from tx_misra.c. Every trace event is stamped zero.
The buffer carries no timing, and the kernel's own check for an entry having
been overwritten -- time_stamp against the entry's own stamp, in the block and
byte allocates and in the system suspend and resume -- compares zero with zero,
so it never fires and a service patches whatever now occupies the slot.

Both ports now read in MISRA builds the clock they already read otherwise,
_tx_linux_time_stamp.tv_nsec, which TX_TRACE_PORT_EXTENSION refreshes on every
recorded event in both forms of the insert. The non-SMP port's non-MISRA macro
carried a trailing semicolon, which made it a statement and is why the MISRA
insert -- which takes the time source as a function argument -- could not use
it; that is dropped and the two branches become one definition. Both headers
keep the _tx_misra_time_stamp_get declaration, because tx_misra.c still defines
it and is compiled for these ports.

The MISRA insert evaluates its time source before the callee refreshes the
clock, so each entry carries the reading taken at the previous recorded event.
Stamps are real, distinct and ordered, which is what the overwrite check needs.

The trace entry update test gains an assertion that the buffer holds an entry
the port actually stamped. It fails on dev with ERROR #13 under
misra_trace_build and passes with this change. Suites green: 7/7 ThreadX
configurations, 5/5 SMP.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-16 14:01:07 -04:00
Frédéric Desbiens e9b5aec65e Added MISRA build configurations to the ThreadX regression matrix (#747)
TX_MISRA_ENABLE routes the kernel's pointer conversions, its memset and its
stack check through the shim in common/src/tx_misra.c. No build configuration
defined it, so that file was never compiled or tested, and the source the MISRA
analysis runs against was not the source the suite exercises. Adding event
tracing alongside it reaches a further region of the shim that neither macro
alone compiles. Both defects found in this area were invisible for the same
reason: #741 broke under TX_MISRA_ENABLE and #745 under both macros together,
and neither combination was built anywhere.

misra_build and misra_trace_build fill that gap. Two things had to give way for
them. TX_POINTER_TO_ALIGN_TYPE_CONVERT exists only when TX_MISRA_ENABLE is
absent, so threadx_test_port.h writes the conversion out rather than taking it
from the API; it is test scaffolding storing a pointer in a word wide enough to
hold one, not kernel source. And thread_transition and module_manager compile
hand-picked common/src sources directly instead of linking the library, which is
what lets them reach feature macros the library-per-configuration model cannot;
under TX_MISRA_ENABLE those sources call into the shim, and pulling the shim in
pulls its own callees after it. Neither directory exists to test the shim, so the
MISRA configurations skip them.

100/100 tests pass in each of the two new configurations, against 105/105 in the
existing five - the five not run are the four thread transition tests and the
module manager test, exactly the two directories skipped. Build is 4 to 5
seconds and the suites 13 and 15 seconds, against 4 seconds and 9 to 17 for the
configurations already there, so the tx job grows by roughly forty seconds.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-16 09:55:24 -04:00
Frédéric Desbiens 24d1f278c2 Decorrelated the wait abort ISR test's interrupt source from the thread it samples (#743)
threadx_thread_wait_abort_and_isr_test waits for the periodic timer interrupt to
land while the preempt disable flag is set. The same interrupt also puts the
semaphore that wakes the thread being sampled, so the two run in lockstep: the tick
wakes thread 0, thread 0 does a fixed amount of work and suspends, and the next tick
arrives a fixed interval later with thread 0 in the same place every time. That is
the resonance the handler's own comment describes, and perturbing the handler's
duration only shifts the phase rather than breaking the correlation. The window
itself is a few instructions wide, so a sampler locked to the thread's own cycle can
miss it indefinitely, which is what #644 and #649 measured and worked around.

The simulator's timer thread waits on _tx_linux_timer_semaphore with a one-tick
deadline and delivers an interrupt early when the semaphore is posted, which the port
already relies on in _tx_thread_schedule. The test now runs a plain POSIX thread that
posts it, injecting interrupts at moments unrelated to the tick grid and sampling
thread 0 at arbitrary points in its cycle rather than the same one. The injector posts
only when nothing is outstanding, so interrupts can never be queued faster than they
are serviced. It is confined to the Linux simulation port and no port file changes.

The count of windows asked for goes back to ten, on the same reasoning that lowered
it to three: ask for what a run can actually reach. Measured over six build
configurations of a branch carrying this change, every one reaches ten of ten in under
a second, against one to three of three previously with three of the six pinned at the
180 second budget. The budget and the zero window ceiling are untouched, so a run that
somehow still falls behind behaves exactly as it does today.

The SMP copy deliberately does not get the injector, and its count stays at twenty. It
has never been in the slow mode, and injecting interrupts there measurably hurts:
8.9 seconds against 0 to 1 for the same twenty windows, which is the extra interrupt
traffic contending across the simulated cores. Only the tx copy has the problem this
solves, so only the tx copy changes; the budget and ceiling logic stays identical
between them.

Verified on this branch: the tx suite passes 105 of 105 with the test at 0.44 seconds,
and four direct runs reach ten of ten in 0 to 1 seconds. The SMP suite passes 117 of
117 unmodified, with its own test at 0 to 1 seconds over four runs.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-16 09:18:41 -04:00
Frédéric Desbiens b93d1ee92f Fixed the simulator ports and thread create paths so they compile when TX_MISRA_ENABLE is defined (#742)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
r52_fvp / r52 (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
* Fixed the simulator ports so they compile when TX_MISRA_ENABLE is defined

_tx_thread_stack_build() in the four simulator ports converts the fake stack
pointer through TX_POINTER_TO_ALIGN_TYPE_CONVERT and
TX_ALIGN_TYPE_TO_POINTER_CONVERT. tx_api.h defines both macros only in the
non-MISRA branch of its #ifdef TX_MISRA_ENABLE, so with that macro defined the two
names are undeclared and none of the four files compiles.

Reproduced with:

    gcc -m32 -c -DTX_MISRA_ENABLE -I common/inc -I ports/linux/gnu/inc \
        ports/linux/gnu/src/tx_thread_stack_build.c -o /dev/null

which reports both names as implicit declarations and then an int to pointer
assignment. The same command without the define compiles cleanly.

No build configuration under test/tx/cmake or test/smp/cmake defines
TX_MISRA_ENABLE, so CI never compiles these files in that mode. It surfaced on a
branch that carries such a configuration.

The conversions are now written inline, which is what the non-MISRA macros expand
to and what the surrounding port code already does, including the line this
replaced.

The alternative would be the idiom common/src/tx_thread_create.c uses for the same
conversion: an explicit #ifdef selecting the ULONG pair under MISRA. That is not
equivalent here. On __x86_64__ this port defines ULONG as unsigned int and
ALIGN_TYPE as unsigned long long, so a pointer round-tripped through the ULONG pair
loses its top 32 bits. Writing the conversion inline keeps one form that is correct
in both modes and on both widths.

Verified by compiling the Linux port with and without TX_MISRA_ENABLE, both clean.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Fixed the same MISRA build break in the SMP and module manager thread create

_tx_thread_create() in common_smp and _txm_module_manager_thread_create() convert the
thread's stack start through TX_POINTER_TO_ALIGN_TYPE_CONVERT and
TX_ALIGN_TYPE_TO_POINTER_CONVERT with no conditional at all. tx_api.h defines both
macros only in the non-MISRA branch, so with TX_MISRA_ENABLE and
TX_ENABLE_STACK_CHECKING both defined neither file compiles. It is the same defect as
the simulator ports in the previous commit, in two more files.

Reproduced with:

    gcc -c -DTX_MISRA_ENABLE -DTX_ENABLE_STACK_CHECKING \
        -I common_smp/inc -I ports_smp/linux/gnu/inc \
        common_smp/src/tx_thread_create.c -o /dev/null

which reports both names as implicit declarations. Both files now compile with and
without TX_MISRA_ENABLE.

The conversions are written inline for the same reason as the ports: the #ifdef idiom
that common/src/tx_thread_create.c uses selects the ULONG pair under MISRA, which
truncates a pointer wherever ALIGN_TYPE is wider than ULONG. That is reported
separately.

Verified: the SMP regression suite passes 117 of 117, and the ThreadX suite still
builds.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-15 23:19:32 -04:00
Frédéric Desbiens 26e4aa182c Fixed the build of tx_misra.c with MISRA and event trace enabled (#746)
tx_misra.c defines five pointer conversions for the trace component, and their
prototypes sit in tx_trace.h behind TX_SOURCE_CODE. tx_misra.c is the one kernel
source that does not define that macro, so it defines all five with no previous
declaration. Under -Wmissing-declarations -Werror, which the kernel target is
built with, the file does not compile at all, so TX_MISRA_ENABLE and
TX_ENABLE_EVENT_TRACE cannot be enabled together. No build configuration defines
both, which is why this has never surfaced.

The five prototypes move out of the TX_SOURCE_CODE guard into their own block,
next to where tx_api.h already declares _tx_misra_trace_event_insert - the sixth
function of the same region, which escapes the problem for exactly that reason.
Declarations only; no definition moves and no macro changes meaning.

All 185 files in common/src now compile clean under all four combinations of the
two macros with -Wall -Wextra -Wmissing-declarations -Werror, against five errors
in tx_misra.c for the both-enabled case before. Object code is byte identical for
the three combinations that already built: 0 of 185 files differ in each.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-15 23:08:23 -04:00
Frédéric Desbiens 9e4c57138d Normalized the AI disclosure comment to one fixed line per file (#740)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
r52_fvp / r52 (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The per-edit disclosure named the product and model, so every agent and every
model version appended another line. 74 files carried two to four of them, and
the same five products had accumulated 13 spellings -- Copilot against GitHub
Copilot, Claude Sonnet 4.6 against claude-sonnet-4.6, four spellings of Codex.
Twenty assembly lines carried a doubled comment marker, `; //` or `@ //`.

Every file now carries exactly one line, fixed text naming no product:

    Portions of this file were generated with AI assistance.

It is written with the comment character that file already uses, so the `;`
and `@` assembly files keep theirs and the doubled markers are gone. Precise
attribution stays on the commit, where the Assisted-by trailer is per-change,
dated and attached to the diff it describes. A header line cannot hold that
record honestly, because the code it names gets rewritten and the line stays.
A file-level flag answers whether; the history answers who.

Comment-only. 455 files, 455 insertions and 574 deletions: every removed line
was a disclosure line, every added line is the fixed text, and no file is left
with zero or with more than one. `scripts/check_ports.sh` passes, including the
reproducibility check that would catch a ports_arch master and its generated
copies drifting apart. Recompiled against dev, every file that builds without a
vendor toolchain gives a byte-identical object: 19 of 19 C files under common,
100 of 100 GNU assembly files, and all 16 assemblable files whose comment
marker changed. The 10 remaining marker changes are ac5 and IAR sources where
`;` already started the comment and only the redundant `//` was removed.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-15 17:17:26 -04:00
Frédéric Desbiens f5e59a1dd3 Completed Windows simulator support and regression coverage (#736)
Completed and stabilized the Win32 and Win64 MSVC simulator ports.

- Replaced high-latency host synchronization with bounded scheduler handoffs and
  critical sections.
- Added a high-resolution timer, tick batching, idle fast-forward, shutdown
  coordination and Win64 extension-pointer support.
- Brought the Windows regression tooling and the new thread-transition tests up
  on CMake, Ninja and the Visual Studio Build Tools.
- Extended the SMP teardown diagnostics and corrected 64-bit trace-test handling.

This supersedes the historical `win64`, `win32-perf`, `windows-sim-ports` and
`windows-sim-ports-completion` branches; no unmerged change from them is missing
here. The original Win64 port landed in #529.

1,610 of 1,610 tests pass, across five configurations each: Win32 515, Win64 515,
Win64 SMP 580. No external dependency was added, and the existing MSVC warnings
in the trace configuration are unchanged. Hardware validation does not apply to
host simulator ports.

Assisted-by: Codex (gpt-5.6-sol) <codex@openai.com>
2026-09-15 16:47:08 -04:00
Frédéric Desbiens 30b3d22adc Restored the random stack fill value cleared during thread creation (#732)
Fixes #723

With `TX_ENABLE_STACK_CHECKING` and `TX_ENABLE_RANDOM_NUMBER_STACK_FILLING` both
enabled, `_tx_thread_create` picked a random byte, stored it in
`tx_thread_stack_fill_value`, filled the stack with it -- and then cleared the
whole control block with `TX_MEMSET`. The stack held the pattern while the
control block claimed zero, so `TX_THREAD_STACK_CHECK` and
`_tx_thread_stack_analyze` compared against the wrong value for the entire life
of the thread.

The value is now computed into a local and written back after the clear. The
clear stays where it is, because the module manager's error checking walks the
created list before the control block may be touched.

Fixed in `common/src/tx_thread_create.c`, `common_smp/src/tx_thread_create.c`
and `txm_module_manager_thread_create.c`, where it additionally left a user-mode
module thread's kernel stack filled with zeros.

`threadx_thread_stack_fill_value_test` creates sixteen unstarted threads and
checks each control block against the pattern in its stack, tolerating a random
zero byte without letting that hide the defect. Added to both suites: `ERROR #3`
on `dev` in `stack_checking_rand_fill_build`, green with the fix. Five tx
configurations at 104 tests, five SMP at 117, and all three sources clean under
`-Wall -Wextra` across every combination of the four stack-filling switches.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-15 16:24:58 -04:00
Frédéric Desbiens 201cc04609 Applied R_ARM_RELATIVE relocations at GNU ThreadX module startup (#731)
Fixes #230

`gcc_setup.s` rebases the GOT and copies initialized data into the module data
area, but never applies the relocations the linker records for words inside that
data. A global initialized with a link-time constant -- `char *pBuffer = buffer;`,
a function pointer, a struct member holding a string literal's address -- kept
its link address and pointed outside the loaded module.

`gcc_setup.s` now walks `.rel.dyn` after the data copy and adjusts every
`R_ARM_RELATIVE` entry targeting the module data area, reusing the GOT loop's
base arithmetic and skip rules, and reloading the flash base first because
`crt0_memory_copy` clobbers `r3`.

The table is only emitted with `-pie`, so the example build scripts now pass
`-pie --no-dynamic-linker` and the example linker scripts discard `.dynamic`,
which `-pie` would otherwise place at address zero and inflate the image to
nearly 200 KB. Without `-pie` the table is empty and the loop is a no-op.

Covers Cortex-M3, M4, M7, M33 and M0+. Cortex-M23 ships no module linker or
build script and AArch64 uses a different relocation format; both left for a
follow-up.

Verified on `qemu-system-arm` with a module linked at one address and loaded at
another: char, function, string and struct-member pointers all resolve to the
loaded addresses, and the same module built without `-pie` still runs.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-15 16:24:54 -04:00
Frédéric Desbiens 3e5a16c95b Allowed tx_timer_change to be called from tx_application_define (#730)
Fixes #224

`_txe_timer_change` returned `TX_CALLER_ERROR` for any call at or above
`TX_INITIALIZE_IN_PROGRESS`, making `tx_timer_change` the only timer service
that could not be called from `tx_application_define` -- while still being
allowed from an ISR.

The check has no technical basis: `_tx_timer_change` only writes the expiration
fields of a timer that is not on an active list, with interrupts disabled.
Removed it from both the `common` and `common_smp` copies of
`txe_timer_change.c`, with the now-unused includes and the `TX_CALLER_ERROR`
line in the header comment. Relaxing an error check is backward compatible.

`testcontrol.c` in both suites now calls `tx_timer_change` at initialization, so
the timer simple test covers it: `ERROR #30` on `dev`, green with the fix.
116/116 SMP, 103/103 non-SMP.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-15 16:24:50 -04:00
Frédéric Desbiens 9b2979e6b0 Replaced the GNU-only dsb/isb 0xF operands with the UAL sy form (#729)
Fixes #551

`_tx_thread_system_return_inline()` in the Cortex-M `tx_port.h` headers spells
its barriers `dsb 0xF` and `isb 0xF`. A bare hexadecimal operand is a GNU
assembler extension, and IAR rejects it with `operand syntax error`, so the
header cannot be included at all. The block is guarded for GCC, armclang and IAR
together, so every IAR user of an affected port hits it -- four independent
reports on Cortex-M33 and M7 with EWARM 9.50 and 9.70.

Both operands become `sy`, the Arm UAL name for exactly what `0xF` encodes. The
generated instruction is unchanged. Applied to the two `ports_arch` masters and
all 32 copies under `ports`, covering M0, M23, M3, M33, M4, M52, M55, M7 and M85
across ac5, ac6, gnu, iar and keil, plus the `scripts/check_ports.sh` probes that
matched the old spelling.

`check_ports.sh` passes, the copy scripts still reproduce every generated port
byte for byte, and `arm-none-eabi-gcc -O2` compiles a caller for every patched
header, emitting `dsb sy` and `isb sy`. Three headers that need toolchain
intrinsics GCC does not ship fail identically on `dev`.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-15 16:24:45 -04:00
Frédéric Desbiens d15f28ab9f Masked the SMP remap core maps to silence a false -O2 array bounds error (#728)
Fixes #469

`_tx_thread_smp_remap_solution_find` indexes `_tx_thread_smp_schedule_list` with
the lowest set bit of the core maps it is given. Every caller already masks
those maps with `TX_THREAD_SMP_CORE_MASK`, but the compiler cannot see it, so
when the function is inlined into `_tx_thread_system_suspend` at `-O2` GCC
assumes a bit number as high as 31 and reports an out-of-bounds subscript.
`-Werror` turns that into a build failure.

Masked the three incoming maps at the top of the function, in both the inline
version in `common_smp/inc/tx_thread.h` and the twin in
`tx_thread_smp_utilities.c`. The masks are semantic no-ops, so scheduling is
unchanged, but the range is now visible to the optimizer. The first core queue
entry is also initialized, because the narrowed range lets GCC consider an empty
queue and warn about that instead.

Reproduced on the Cortex-A9 SMP port with arm-none-eabi-gcc 13.2.1, and clean
afterwards across -O2, -O3 and -Os and 2, 4 and 8 core configurations. SMP suite
116/116.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-15 16:24:40 -04:00
Frédéric Desbiens 1d4a4aeec8 Guarded _tx_thread_stack_analyze against inverted stack pointers (#727)
Fixes #460

`TX_ULONG_POINTER_DIF` casts the pointer difference to `ULONG`, so when
`tx_thread_stack_highest_ptr` sits below `tx_thread_stack_start` the midpoint
wraps to a huge value, the probe lands outside the stack, and the search never
converges: the caller hangs or faults.

`_tx_thread_stack_analyze` now requires the highest pointer to be strictly above
the start of the stack, and bounds the final scan by it. #464 covered the
`TX_THREAD_STACK_CHECK` path; this covers direct callers too, as
@billlamiework suggested on the issue.

Inconsistent pointers still mean a real overflow or a corrupted control block,
which remains the application's problem. What changes is that ThreadX reports it
through the stack error handler instead of hanging.

New cases for an inverted and an equal pointer pair crash the suite without the
fix and pass with it.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-15 16:24:35 -04:00
Frédéric Desbiens 5d235a534c Fixed the garbage _tx_initialize_unused_memory in the GNU Cortex-A ports (#726)
Fixes #435

`LDR x1, =__top_of_ram` already loads the top of RAM, so the `LDR x1, [x1]`
that followed read whatever sat at that address and left
`_tx_initialize_unused_memory` holding garbage.

Dropped that instruction from the 13 non-SMP GNU Cortex-A ports, the ARMv8-A
source they are generated from, and the two Cortex-A35 module examples. The SMP
GNU ports already had the correct form, and the Arm Compiler ports are
unaffected -- their symbol really does need the dereference.

`scripts/check_ports.sh` passes and the `gnu` CI job build-verifies the AArch64
ports. Not run on hardware.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-15 16:24:09 -04:00
Frédéric Desbiens cd6a2d9034 Put the module manager test's thread on the kernel's created list, so the stand-in kernel matches the one the manager reads (#739)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The host test built a TX_THREAD, gave it an ID and a module instance, and
left it off the kernel's created list. A created thread lives on that list,
so the thread the test created was not one the kernel would recognise.

The created list head and count are now defined for each object type, the
thread the test creates goes onto the thread list the way thread create puts
it there, and a successful delete takes it off again the way thread delete
does. Only the thread list is populated; the others stay empty because
nothing here creates an object of those types.

No expectation changed.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-14 17:09:00 -04:00
Frédéric Desbiens 83a3e62c28 Added a Module Manager host test harness and released a module thread's kernel stack only when it has one (#738)
* Released a module thread's kernel stack only when it has one, so deleting a thread without a manager-allocated stack no longer asks for that memory back

The thread delete dispatcher decided whether to release a kernel stack from
the calling module's property flags. A user-mode module can be given the
address of a thread that carries no kernel stack the manager allocated, and
deleting it passed that thread's null stack pointer to
_txm_module_manager_object_deallocate().

The pointer is now tested instead of the module's properties. Both thread
create paths clear the whole control block before filling it in, so a thread
with no manager-allocated kernel stack holds TX_NULL in the field, which
makes the test exact rather than an inference from how the module was
loaded.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Added a Module Manager host test harness and covered the user-mode thread kernel stack lifetime

The Module Manager is not in the ThreadX library the test tree links, and its
control blocks come from a module port rather than a base port, so nothing in
the suite could reach it. The thread_transition directory already describes
this technique as the one "the module manager tests in this tree use", but no
such tests existed.

This adds the directory that comment refers to: a port shim that replaces the
three interrupt primitives with host equivalents that count, and a CMake
target that compiles the manager sources under test directly against one
module port's headers.

The dispatch layer is one header of static functions and an unoptimised build
emits all of them, so a test that includes it to reach one dispatcher would
pull in references to every service the manager can dispatch. The header
guards each dispatcher with its own TXM_<SERVICE>_CALL_NOT_USED macro, so the
build reads that list out of the header and defines every guard except the
ones the test needs. Reading it rather than writing it down means a service
added later is excluded without anyone having to remember.

The first test asserts the invariant a user-mode module thread has to hold:
one logical thread costs the object pool two allocations, the control block
the module asks for and the kernel stack the manager takes on its behalf, and
deleting the thread must return the pool's available bytes and the module's
allocation-list count to exactly what they were. It asserts an equality
rather than a bound, because a bound would pass while one of the two was left
behind on every cycle, and then runs enough cycles that a leak of one stack
per cycle exhausts the pool several times over.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-14 16:08:48 -04:00
Frédéric Desbiens dde43b8ab2 Replaced the stale system stack switch pseudo-code in the ARMv7-A ports with comments that describe what the code actually does (#735)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The context save, vectored context save and system return routines in the
ARMv7-A ports carried pseudo-code comments claiming that they saved the
thread stack pointer and then switched to _tx_thread_system_stack_ptr.
Neither of those things happens, and none of these ports references that
variable outside an unused IMPORT in their example builds.

On ARMv7-A each processor mode has its own banked stack pointer. The IRQ
handler branches to _tx_thread_context_save while still in IRQ mode, so the
core's banked IRQ stack already serves as the system stack, and the thread
stack pointer is stored in the control block by _tx_thread_context_restore,
and only when the interrupt results in preemption. The scheduler runs on the
banked SVC mode stack that the startup code sets up. There is nothing for a
software stack switch to do.

The comments were therefore misleading rather than merely redundant, and had
led at least one user to try to restore the code they described. They are now
replaced by a description of the actual mechanism.

The AArch64 SMP ports keep their comments unchanged, because ARMv8-A does not
bank a stack pointer per processor mode and those ports do reload
_tx_thread_system_stack_ptr[core] explicitly.

This is a comment-only change. Every changed line is a comment, and all
twenty-five GNU variants still assemble cleanly for their target core.

The fourteen files under ports/cortex_a{5,7,8,9,12,15,17} were regenerated
from ports_arch/ARMv7-A/threadx/common/src/tx_thread_system_return.S with
ports_arch/ARMv7-A/update.sh. The ARMv7-A SMP ports have no generator, so
those files were edited directly.

Fixes #734

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-11 13:27:32 -04:00
Frédéric Desbiens c0aa4dbe29 Initialized the suspend status before the non-interruptable suspend in tx_thread_sleep, so a sleep no longer returns a stale error left over from an earlier timed-out suspension (#725)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
r52_fvp / r52 (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
When TX_NOT_INTERRUPTABLE is defined, _tx_thread_sleep did not set
tx_thread_suspend_status to TX_SUCCESS before calling
_tx_thread_system_ni_suspend. The interruptable path always did.

Nothing else writes that field on behalf of a sleeping thread:
_tx_thread_timeout resumes a TX_SLEEP thread directly and there is no
suspend cleanup routine for sleep. The value therefore survived from
whatever suspension the thread performed last, and tx_thread_sleep
returned it. A tx_semaphore_get that timed out with TX_NO_INSTANCE
would be followed by a tx_thread_sleep that also returned
TX_NO_INSTANCE after sleeping correctly.

The same change is applied to the SMP copy, which is identical.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-10 12:06:54 -04:00
Frédéric Desbiens 70741198e9 Refused thread delete and reset while an exit transition is in progress (#724)
_tx_thread_shell_entry and _tx_thread_terminate both publish a thread's
terminal state -- TX_COMPLETED or TX_TERMINATED -- and then call that
thread's exit notification callback, before the thread has been detached
from the ready list and before either service has finished with the pointer
it holds to the control block. That terminal state is exactly the state
_tx_thread_delete and _tx_thread_reset accept as authorization to
invalidate or rebuild the control block, and neither service tested whether
the transition producing it had finished.

A callback could therefore delete the thread it was called for -- and then
lawfully recreate it over the same memory, since delete exists to permit
that -- while the scheduler was still linked to the old incarnation.
tx_thread_create zeroes the whole control block and can auto-start the new
one, so the old priority list is left heading at a block whose own priority
field names a different list, with the old priority's map bit set behind
nothing. Alternatively a callback could reset a terminated thread, which
moves it out of the terminal state, and then resume it: the interrupted-
suspension logic in _tx_thread_system_resume refuses to void a suspension
only while the state is still terminal, so with the reset allowed first the
resume clears the suspending flag and restores TX_READY, the outer service
then finds the flag clear and skips the removal, and tx_thread_terminate
returns TX_SUCCESS for a thread that is runnable again.

Interrupt masking does not close the window, because the kernel restores
the prior posture before invoking the callback deliberately. On SMP
TX_RESTORE also releases the global protection, so the target can be
executing on another core while its callback runs -- and a reset there
memsets the stack a live core is running on. On the Linux, Win32 and Win64
host simulation ports the consequence is more immediate than corruption:
TX_THREAD_DELETE_PORT_COMPLETION cancels and joins the host thread backing
the deleted thread, so a callback-side delete of the completing thread
destroys the host thread the callback is running on.

The fix marks the transition and has the two services refuse a marked
target, returning the errors they already document, TX_DELETE_ERROR and
TX_NOT_DONE. The refusal is transient and the same call succeeds once the
transition has completed, so no documented lifecycle is lost; and it is in
the core services rather than the _txe_ wrappers, so disabling error
checking cannot disable it. Refusing the reset is also what closes the
resume path, without touching _tx_thread_system_resume: its existing
terminal-state test is sufficient once nothing can turn the terminal state
into TX_SUSPENDED from inside the window.

tx_thread_suspending is the marker, rather than a new control-block field.
It already means "a suspension is in progress" and is already true across
the callback in the two interruptable paths, so no field is added, the
public structure is unchanged, and sizeof(TX_THREAD) is unchanged --
which matters, because the Module Manager's object handling depends on the
sizes of the control blocks. Widening its lifetime was checked against
every reader rather than assumed. There are four: two in
_tx_thread_system_suspend and two in _tx_thread_system_resume. In every
window this change widens, the state is TX_COMPLETED or TX_TERMINATED, and
both resume readers already refuse to void a suspension for exactly those
two states, so their behaviour is unchanged; and no suspension routine is
called on the target in those windows, so the suspend readers never see
them. No suspension-initiating service can set the marker again inside a
window either: every one of them acts on a thread that is ready or
suspended.

Three sites needed changing beyond the two refusals, and the shape of each
was decided by where the marker can safely be cleared:

  - The non-ready branch of _tx_thread_terminate cleared the marker before
    the terminated extension and the callback, which is what left them free
    to act on a control block the service still had mutex-release
    processing to do against. The clear moves to the common tail, after the
    last dereference of the target, and becomes the single clear site for
    the whole service. In the interruptable ready branch the flag is
    already false there, because _tx_thread_system_suspend cleared it when
    it detached the thread, so the tail store is a second store of a value
    the flag already holds -- cheaper than testing for it, and it keeps one
    clear site.

  - Under TX_NOT_INTERRUPTABLE neither path set the marker at all, because
    that configuration does not use the interruptable suspension path that
    sets it. Both now set it before the callback. Interrupts being disabled
    there does not help: the callback is reached by a direct call.

  - In the TX_NOT_INTERRUPTABLE completion path the marker is cleared
    before _tx_thread_system_ni_suspend rather than after it. That call
    returns to the scheduler for a thread that is the current thread, which
    a completing thread is, and does not come back; clearing afterwards
    would leave a normally completed thread marked for ever and therefore
    permanently undeletable. Nothing is lost by clearing early there,
    because everything from that point to the detachment runs with
    interrupts disabled and calls no application code.

The change is the same change twice. All four files are byte-for-byte
identical between common and common_smp at this commit and stay so after
it, so common_smp was written by copying rather than by repeating the
edits. tx_thread_system_suspend.c and tx_thread_system_resume.c, which do
differ between the kernels, are deliberately untouched.

Tests. The in-tree regression test goes to both trees and is byte-for-byte
identical between them. It drives seven scenarios: terminating a ready
non-current target with two peers ready at the same priority, with the
callback attempting the delete and recreating the block if it succeeded;
the same with the callback attempting the reset and then the resume;
terminating a target suspended on a semaphore while owning a mutex, which
is the non-ready branch; natural completion alone at its priority,
including the safe post-completion reset, terminate, delete and recreate at
another priority; self termination; a benign callback, whose notification
count and ordering are unchanged; and the state and boundary cases, where
the new refusal must not fire.

It measures rather than describes. The callback records the published
state, the marker, and the status of every lifecycle service it can reach,
calling the core service as well as the wrapper wherever a refusal is
expected. A snapshot taken under interrupt lockout -- which is the global
SMP protection on an SMP port -- checks that every ready list agrees with
the control blocks it heads, that the priority map agrees with the lists,
and that each execute pointer is a member of the list its own priority
field names. Every walk is bounded, so a corrupted ring costs an assertion
and not a hang, and no test in the suite can hang. Expectations are counted
inside a scenario and gated between scenarios, so a failing kernel reports
how much it failed by without being driven further into its own
corruption.

The consequences are demonstrated from the terminator's context rather than
the completing thread's, which is what makes the pre-fix behaviour an
assertion instead of a wedged simulator. Compiled against the unfixed
sources the test fails 9 of the 24 expectations it reaches in the
uniprocessor tree and 10 of 24 in the SMP tree, and the failures are the
finding: the callback-side delete succeeds, the recreate succeeds, the
consistency snapshot disagrees, the target is neither terminal nor detached
when the service returns, and a peer has left the ready ring the recreated
block hijacked.

The SMP tree gets a second test for the case that needs concurrency. The
victim is excluded to core 1 and spins there without relinquishing while
the controller, excluded to core 0, terminates it, so the callback runs on
one core while the target executes on another. The callback-side reset and
delete must both be refused, and a sentinel written into the unused low end
of the victim's stack must survive -- a reset would have memset the whole
stack before rebuilding the frame. Every wait is bounded, and if the remote
precondition cannot be established the test says so and drops only the
assertions that depend on it rather than reporting a pass it did not earn;
measured over twenty consecutive runs it established the precondition every
time.

TX_NOT_INTERRUPTABLE and TX_DISABLE_ERROR_CHECKING are not among the five
build configurations either tree compiles, and each tree builds the whole
library once per configuration, so neither can be reached from inside the
suites. The lines this change adds under TX_NOT_INTERRUPTABLE are therefore
in no configuration the trees build, and they are where the permanent-
undeletability failure mode lives, so they get their own harness rather
than a compile check: the four sources plus the two error wrappers are
compiled directly into a test executable, once per combination, with
recorders standing behind the scheduler services they call. That is what
makes the marker's value at the moment of detachment directly observable.
It runs 66 expectations under TX_NOT_INTERRUPTABLE, 66 under that with
error checking disabled, 45 under that with notification disabled, and 63
under error checking disabled alone; against the unfixed sources those fail
25, 22, 6 and 20 respectively. The harness lives in the uniprocessor tree
only, because the four sources are identical between the kernels and the
shim replaces the very primitive the SMP port differs in, so a second copy
would compile the same text under the same macros. It is deliberately left
out of the coverage instrumentation, since the same source under different
feature macros has a different line set and merging those would confuse the
union rather than add to it.

Results. Both suites pass in all five configurations with GCC 14: 103 of
103 in the uniprocessor tree, up from 98, and 116 of 116 in the SMP tree,
up from 114. Merged line coverage is 100% in the uniprocessor tree and
5172 of 5183 in the SMP tree, whose eleven uncovered lines are the same
eleven that were uncovered before this change and are in tx_byte_pool_search
and tx_thread_smp_utilities; all four changed files are at 100% line
coverage in both trees, and SMP branch coverage rises from 2819 of 3548 to
2831 of 3556. Cross-compiled with arm-none-eabi-gcc at -Wall -Wextra for
Cortex-M4 against common and for Cortex-A7 SMP against common_smp, all four
files produce no diagnostics at all and an identical warning set to before
the change, under -std=gnu99 and -std=c99 alike -- unlike the module ports,
-std=c99 does not fail on these base ports, and even -Wconversion is clean.

No MISRA deviation is required: explicit comparisons to TX_TRUE, existing
ThreadX types, single-entry and single-exit control flow, no goto, and two
added constant-time tests that change no real-time complexity.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-10 12:04:40 -04:00
Frédéric Desbiens 9ee198a2d6 Made the stack analyze binary search require several consecutive fill words, so unwritten holes in a used stack no longer cause the highest stack pointer to be under-reported (#720)
The binary search in _tx_thread_stack_analyze() accepted a probe location as
unused as soon as a single word still held TX_STACK_FILL. A word inside an
otherwise used region that simply was never written - the padding of a
partially initialized local array, for example - therefore made the search
move away from the real boundary and report far less stack usage than the
thread had actually consumed, which in turn kept the stack guard from firing.

The probe now walks down from the candidate location and requires
TX_THREAD_STACK_ANALYZE_FILL_WORDS consecutive fill words before it treats
the location as unused, stopping early at the lowest location already known
to hold the fill pattern. The new macro defaults to eight words and can be
overridden in tx_port.h; setting it to one restores the previous behavior.
Holes shorter than the configured run no longer mislead the search, and the
result is unchanged for stacks that contain no such holes.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-10 11:02:42 -04:00
Frédéric Desbiens d56f96f37f Refreshed the reason gcc_check leaves RISC-V out (#722)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
r52_fvp / r52 (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The header explained the exclusion by saying RISC-V "is not regressing" and
that adding it would widen the toolchain download. The second half is still
true; the first read as though nothing in CI exercised the family at all,
which stopped being the case with #717.

RISC-V is now the best-covered of the four excluded families rather than the
least: regression_test.yml builds both ports and runs 955 tests on them under
QEMU -- 475 on RV32 and 480 on RV64, across five build configurations each --
which is more than a compile-and-link check could establish. That is a
stronger argument for leaving it out of this workflow than the original, so
the sentence now makes it.

The download figure is kept and quantified: the two bare-metal toolchains this
workflow would have to fetch are about 500 MB apiece.

Comment only; no behaviour change. scripts/check_gcc.sh names the same four
families but states the exclusion without giving a reason for it, so it needs
no matching edit.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-10 09:31:58 -04:00
Frédéric Desbiens 6c84e61d19 Gave install_riscv.sh the network hardening install.sh already had, and shared it between them (#721)
install.sh grew a retry loop, per-command timeouts and a deliberately
non-gating apt-get update after this runner pool cost several whole runs: a
mirror going silent for two hours, and a Hash Sum mismatch from a third-party
repository the project does not even use turning builds red. Those lessons were
local to that one file.

install_riscv.sh had none of them, and #717 puts it on every pull request's
critical path. Under set -e its bare apt-get update was a single point of
failure for the whole suite -- the precise case install.sh downgrades to a
warning on purpose -- and its two wget calls, each fetching about 500 MB, had
no retry and no timeout.

Rather than copy the helpers and let them drift again, they move to
tx_ci_common.sh and both scripts source it, following the arrangement
scripts/tx_windows_common.ps1 already uses on the Windows side. install.sh
keeps its behaviour exactly: same APT_OPTIONS, same 120-second TIMEOUT, same
three-attempt retry, and the comments explaining each of them travel with the
code they explain.

Two things are new:

  - TIMEOUT_LONG, 180 seconds, for a single large download. Sized against the
    39 seconds each tarball took on 10 Sep 2026 and deliberately not larger:
    the install step is capped at ten minutes, and a per-attempt timeout able
    to swallow that cap would leave the retry loop no turn to take, which is
    the failure mode the apt comment already records.

  - fetch(), which verifies a SHA-256 before anything is unpacked. Both digests
    were taken from the releases API and then checked against the bytes the CDN
    actually serves. This is not an independent trust root -- expected value and
    file come from the same host -- but it pins the bytes, so a deleted and
    re-pushed tag or a replaced asset stops the build instead of being picked up
    silently.

Also verifies qemu-system-riscv32 alongside riscv64. run.sh selects one per
architecture, so both are worth failing on here rather than at the first test.

Verified locally: retry returns 0 on success and 1 after three attempts;
fetch accepts a correct digest and, on a wrong one, fails and removes the
partial file; the source line resolves from the repository root, from an
absolute path and through a symlink; and both recorded digests match the
bytes served for the pinned tag.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-10 09:28:44 -04:00
Frédéric Desbiens 42f1d2bcf0 Pointed the SMP coverage comment at the pull request that made its measurements (#719)
The paragraph explaining why the SMP floor sits at 99 rather than 100 ended
with a bare reference that resolves to nothing a reader of this repository can
open. It named a source outside the tree in place of one inside it.

The real source is #677: it wrote the four tests that closed 53 of the 64
lines, took the measurements the paragraph quotes -- the 180,003 windows with
zero handovers, and the three remaining lines in tx_thread_smp_utilities.c --
and raised the floor from 98 to 99. Citing it gives the next reader somewhere
to go.

This is the same correction #718 made across five other files; this line was
missed because it sits in regression_test.yml rather than the template.

Comment only; no behaviour change.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-10 09:06:13 -04:00
Frédéric Desbiens 39277cf026 Stated the compiler and coverage requirements directly, and said who the pinned toolchain default serves (#718)
* Stated the compiler and coverage requirements instead of citing a file no contributor can open

Seven comments across five files cited a maintainer-local document as the
source for two project requirements: that GCC 14 on Linux is the default
compiler, and that the coverage target is 100%. That document is not part of
this repository and is not published anywhere, so the citation gave a reader
nothing to follow -- it named a source they cannot open, in place of simply
stating the requirement.

Both requirements are real and both stay. Only the pointer goes: each comment
now states the requirement on its own terms, which is what the surrounding
prose was already doing everywhere else.

  cmake/cortex_r52.cmake                     the pinned reference toolchain
  scripts/check_gcc.sh                       why the script exists
  .github/workflows/gcc_check.yml            why the workflow exists, and the
                                             GCC_VERSION pin
  .github/workflows/r52_fvp.yml              the GCC_VERSION pin
  .github/workflows/regression_template.yml  the coverage floor, twice

Comments only; no behaviour changes. Two paragraphs are rewrapped where the
shorter text left a ragged line. Verified that scripts/check_gcc.sh still
parses and prints its help from the header range it slices, and that
cmake/cortex_r52.cmake still configures the Cortex-R52 build.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Said who the pinned toolchain default in the CMake toolchain files serves

Both cmake/cortex_r52.cmake and cmake/cortex_m52.cmake default
ARM_TOOLCHAIN_PATH to a toolchains directory under the user's home, guarded by
an EXISTS check. Nothing said whether CI relies on that, and the natural
reading is that it does.

It does not. The three workflows that install a toolchain unpack it into the
workspace and cache it there, and r52_fvp.yml puts that directory on PATH
before configuring; scripts/check_gcc.sh passes -DARM_TOOLCHAIN_PATH at each of
its three CMake call sites. On a runner the guarded directory is absent, the
EXISTS check falls through, and the compiler comes from PATH. The default only
ever fires on a developer machine, where it is what makes a no-flag build work.

Both comments now say that, so the default is not mistaken for a CI dependency
and not removed as dead code. cortex_r52.cmake carries the explanation and
cortex_m52.cmake refers to it, matching the cross-reference already there.

The r52 comment also claimed absolute paths mean "the build does not depend on
PATH ordering", which is only true where the pinned directory exists -- in CI
the build depends on PATH and nothing else. Qualified accordingly.

Comments only; no behaviour changes. Verified that both toolchain files still
configure, and that the fall-through is real: with HOME pointed at a directory
holding no toolchains, cortex_r52.cmake configures against the arm-none-eabi-gcc
found on PATH.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-10 07:55:11 -04:00
Frédéric Desbiens 5e9d046b12 Compiled the module manager C sources with GCC, as clang already did (#716)
#689 added the module manager stage to check_clang.sh alone. The GCC half was
never written, so the module manager C stayed unbuilt by the project's declared
default compiler: 28 files of portable module manager under common_modules,
plus the three to nine per-port files under
ports_module/<core>/gnu/module_manager/src, across nine Arm module ports.

The stage is deliberately check_clang.sh's, port for port and header for
header, because a port covered by one check and not the other implies a parity
the checks list does not have. The same two details the ports dictate carry
over: an SMP port's control blocks come from common_smp rather than common, and
the TrustZone ports need -mcmse for their cmse_nonsecure_entry functions to be
honoured rather than ignored. The clang-only waiver does not: GCC implements
the optimize attribute that tx_thread_secure_stack.c carries, so nothing needs
suppressing for that file.

One divergence is forced by the toolchain. txm_module_manager_absolute_load.c
carries a #pragma message steering callers to the extended entry point, and the
C stages treat any compiler output as a failure. check_clang.sh silences it with
-Wno-#pragma-messages; GCC has no equivalent, and neither -Wno-pragmas nor any
other -W option suppresses the note -- verified with 14.3.rel1. The note is
therefore filtered out of the stage's output instead, together with the source
quote GCC prints beneath it. The filter stops at the next line that begins a
diagnostic of its own, so an error immediately following a waived note is still
reported; that case is what the injected-defect run below checks.

The workflow needed no trigger change: #689 added common_modules/** to both
path lists in advance, for the stage that had yet to arrive. Its header comment
is brought in line with what the workflow now runs.

Verified with the toolchain CI pins, arm-gnu-toolchain 14.3.rel1, both triples.
The full script passes and every port compiles every file, matching the counts
check_clang.sh reports for the same nine ports:

  cortex_a35 31/31   cortex_a35_smp 31/31   cortex_a7 34/34
  cortex_m0+ 33/33   cortex_m23 37/37       cortex_m3 33/33
  cortex_m33 37/37   cortex_m4 33/33        cortex_m7 33/33

Verified that the stage fails as intended by injecting defects into a throwaway
worktree: one in common_modules on the line straight after the waived pragma
note, reported under all nine ports, and one in a single port's own source,
reported only under that port.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-10 07:25:37 -04:00
Frédéric Desbiens 7f6e496233 Added the missing VFP enable field to the Cortex-R5/AC5 thread control block, which VFP builds were writing over the FileX pointer (#715)
The Cortex-R5/AC5 port defines TX_THREAD_EXTENSION_2 as empty while its
assembly reads and writes the per-thread VFP enable flag at [thread, #144].
With TX_ENABLE_VFP_SUPPORT that offset lands on tx_thread_filex_ptr, so
tx_thread_vfp_enable corrupted the FileX pointer and lazy save/restore
tested an unrelated value.

Defined TX_THREAD_EXTENSION_2 as ULONG tx_thread_vfp_enable, matching the
Armv7-A ports and the Cortex-R4/R5 GNU and AC6 ports fixed earlier, and
added tx_port_offset_check.c so a future layout change breaks the build
instead of silently retargeting the accesses. The offset was measured at
144 for this port and asserted.

Fixes #382

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-09 17:17:37 -04:00
Frédéric Desbiens 383dd311cd Removed the vector table offset register and system stack pointer setup from the Cortex-M low-level initialization (#714)
_tx_initialize_low_level wrote VTOR and derived _tx_thread_system_stack_ptr from
the reset vector on every Cortex-M port. Programming VTOR is the job of the
low-level startup code that runs before the kernel is entered, and doing it again
inside ThreadX silently overrode a vector table that the application, a
bootloader or a firmware update had already installed. It also forced every
application to export a vector table symbol that the kernel itself never needed.

_tx_thread_system_stack_ptr is never read by any Cortex-M port, so the value
copied out of the reset vector served no purpose.

Both blocks are removed from all Cortex-M ports, together with the now-unused
symbol declarations. The GreenHills files keep the NVIC base address load that
the removed block used to leave in r0 for the SysTick setup that follows.

Applications that relied on ThreadX programming VTOR must now set it in their
startup code. The bundled examples link their vector table at address zero and
run on the reset default.

Fixes #370

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-09 16:28:59 -04:00
Frédéric Desbiensandr 16297d7a2a Added a workflow that runs the Cortex-R52 images on the Armv8-R AEM FVP (#688)
* Added a workflow that runs the Cortex-R52 images on the Armv8-R AEM FVP

Nothing in CI executed a single instruction of any ThreadX port. gcc_check
compiles and links -- its own header says it "executes nothing" -- and its
CMake stage covers one R52 configuration, the base FVP example. A change that
assembles cleanly, links cleanly and then hangs on the first context switch
passed every required check.

This workflow builds both supported configurations and runs them on the
Armv8-R AEM FVP, asserting each image's self-reported result. The feature
configuration matters as much as the default one: VFP with the hard float ABI,
FIQ, IRQ nesting and FIQ nesting gate whole assembly blocks that the default
configuration never assembles, and it builds eight images where the default
builds five.

Arm distributes the model free of charge but behind a click-through licence
with no stable unauthenticated URL -- the developer.arm.com permalink forms
all 404 for it -- so the download location is a repository variable rather
than a literal. When it is unset the build lanes still run and still gate the
pull request; only execution is skipped, and it says so in the log and in the
job summary rather than passing quietly.

The image list is read from the generated ninja graph rather than kept in the
workflow. The images are EXCLUDE_FROM_ALL, so a bare build reports "no work to
do", and reading the graph means a target added to CMakeLists.txt cannot
escape the check by nobody remembering to list it here.

The module manager port and the S32Z280 targets are not on dev yet. Their
lanes belong in this workflow when they land, not in a second one.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Kept the FVP cache honest and stopped ctest passing on an empty suite

Two corrections to the new workflow.

The FVP cache key named the URL and hashed the workflow file. hashFiles()
reads files, so it cannot hash a repository variable, and pointing it at
this workflow keyed the cache on the file's own contents instead. That is
wrong in both directions: repointing FVP_AEMV8R_URL at a different model
kept serving the old one from cache, which is exactly what the comment
claimed the key prevented, and editing anything in this file forced a
needless refetch of a large archive. A small step hashes the URL and the
key uses that, so it now tracks what it names.

ctest reported success when it found nothing to run. Its default is to
pass an empty suite, and the example registers its tests inside both
if(FVP found) and if(Python3_FOUND). Either one going unmet on a runner
would have left this stage green having executed no image at all, which
is the hole the workflow exists to close. --no-tests=error makes an empty
suite a failure.

Verified with the pinned toolchain, Arm GNU 14.3.rel1: both configurations
configure, enumerate and build, and the image lists are the ones claimed.

  default: 5 images  boot_check demo_m2 demo_m3 demo_mpu demo_threadx
  feature: 8 images  the above plus demo_fiq demo_m5 demo_nesting

Also checked that the target enumeration survives the real ninja graph:
the generated graph offers twenty .elf paths, and the filter reduces them
to those eight, dropping the cmake_object_order_depends_target aliases and
the CMakeFiles copies.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

---------

Co-authored-by: r <r@r>
2026-09-09 14:39:54 -04:00
Frédéric Desbiensandr 8be4c45cce Compiled the module manager C sources, which no check had ever built (#689)
* Compiled the module manager C sources, which no check had ever built

#672 corrected this script's assembly glob and brought the module ports into
the count, but only their assembly. Their C stayed outside every check: 27
files of portable module manager under common_modules, plus the per-port code
under ports_module/<core>/gnu/module_manager/src. By this script's own standard
-- a port simply absent from the count reads as covered -- 293 files across
nine Arm module ports were compiled by nothing, with either compiler.

Each module port ships its own tx_port.h and txm_module_port.h carrying the
control-block extensions the dispatch code needs, so a port is compiled against
its own headers rather than the base port's. Two details the ports themselves
dictate:

  An SMP port's control blocks come from common_smp. Pairing cortex_a35_smp
  with the single-core headers hid _tx_thread_smp_protect and
  _tx_thread_smp_unprotect behind implicit declarations and lost
  tx_thread_smp_core_executing from TX_THREAD, so fourteen files reported
  errors for a port that builds correctly.

  The TrustZone ports carry cmse_nonsecure_entry, which needs -mcmse to be
  honoured rather than ignored. tx_thread_secure_stack.c also carries GCC's
  optimize attribute, which clang does not implement; that divergence is
  suppressed by name, for a file GCC builds cleanly.

All nine ports compile: 293 of 293. Verified that the stage fails as intended
by injecting a defect into a throwaway worktree -- a defect in common_modules
is reported under every port, one in a port's own source only under that port.

Assisted-by: Claude Code (Opus 5)

* Ran the module manager stage against the sources that reach it

Three corrections to the new stage, all found by running it against dev
rather than against the tree it was written on.

The workflows did not trigger on common_modules. Both check_clang.yml and
check_gcc.yml list ports_module but not common_modules, and the portable
module manager under it is the larger half of what the stage compiles --
28 of the 31 to 37 files each port builds, plus two include directories.
A change there left the stage unrun, which is the same "absent from the
count reads as covered" that the stage exists to close. Both lists gain
common_modules, and they stay identical to each other as the comment in
each asks.

A deliberate deprecation notice read as a build failure. Since this PR
was opened, txm_module_manager_absolute_load.c gained a #pragma message
steering callers to the extended entry point. The stage treats any
compiler output as a failure, so that one notice failed every port: nine
failures on a tree where nothing is wrong. Pragma messages are now
waived for the stage, because a notice to callers is not a defect in the
file that carries it.

That failure also printed nothing. Both C stages report by grepping the
output for "error:", so a diagnostic that is not an error produced a bare
FAIL line with no reason under it, and the only way to learn the reason
was to reproduce the compile by hand. Both stages now fall back to
showing what the compiler actually said.

The TrustZone attribute waiver is narrowed to the one file that needs it.
tx_thread_secure_stack.c carries GCC's optimize attribute, which clang
does not implement; it is the only file among the 300-odd this stage
compiles that does. Waiving the warning for the whole port would have
swallowed a stray unknown attribute anywhere else in it.

Verified with the same toolchain CI uses, ATfE 22.1.0: the full script
passes, and every port compiles every file.

  cortex_a35 31/31   cortex_a35_smp 31/31   cortex_a7 34/34
  cortex_m0+ 33/33   cortex_m23 37/37       cortex_m3 33/33
  cortex_m33 37/37   cortex_m4 33/33        cortex_m7 33/33

The counts are each one higher than this PR first reported, because
txm_module_manager_absolute_load_extended.c has landed since. The stage
picked it up with no edit, which is what globbing the directories was for.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

---------

Co-authored-by: r <r@r>
2026-09-09 14:39:30 -04:00
Frédéric Desbiens 6ce8d5cc76 Marked every published ThreadX include directory as SYSTEM so applications no longer get warnings from ThreadX headers (#713)
* Marked every published ThreadX include directory as SYSTEM so applications no longer get warnings from ThreadX headers

Commit 8c3c08f added the SYSTEM keyword to the target_include_directories
call in common/CMakeLists.txt, but the same call in every port, in the
SMP common directory's consumers, in the POSIX and FreeRTOS compatibility
layers, and for the generated tx_user.h directory was left unchanged.
A consumer building with a strict warning set therefore still saw
diagnostics coming from tx_port.h and from the compatibility layer
headers, which is exactly what the original change set out to avoid.

Added the SYSTEM keyword to all of those calls so CMake emits -isystem
rather than -I for every directory holding a ThreadX public header. The
ARMv7-M, ARMv8-M, ARMv7-A and ARMv8-A architecture sources under
ports_arch were updated alongside the ports they generate, keeping the
two in step.

Directories that are PRIVATE to an example or test build were left as
they are, since nothing is published from them.

Fixes #290

Assisted-by: Copilot (Opus 5) <noreply@github.com>

* Stopped apt-get update being a gate it was never meant to be

apt-get update fails if any configured repository serves a bad index,
including ones this project never reads. The GitHub runner image carries
Google's and Microsoft's apt repositories, and a Hash Sum mismatch from
Google's, their CDN caught mid-publish with the index and the Release
file eight hours apart, failed all three attempts and turned a run red
over a browser nobody was installing.

Made a failed update warn and carry on, leaving apt-get install as the
gate. Nothing is weakened by that: the install still exits on a package
it cannot find, so an unreachable archive still stops the script, one
step later and naming the package it could not get, which is a better
diagnostic than a hash mismatch in a repository nobody asked for.

Disabling third-party sources before updating would keep the update
strict, but this script also runs on a contributor's own machine, and
rewriting someone's apt configuration to suit CI would be worse than
tolerating a stale index for an archive we do not read.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-09 14:27:02 -04:00
Frédéric Desbiens 3fc28d8979 Allowed a BASEPRI-masked interrupt to wake the Cortex-M idle loop from WFI (#711)
The idle loop in _tx_thread_schedule masks interrupts before re-reading
_tx_thread_execute_ptr, so that an interrupt cannot make a thread ready
between that read and the WFI and then be lost. With PRIMASK that is safe:
the architecture excludes PRIMASK from the WFI wake-up condition, so the
interrupt still wakes the processor and is taken as soon as the loop
re-enables interrupts.

BASEPRI is not excluded. When TX_PORT_USE_BASEPRI is defined, the mask
written at the top of the loop is still in place at the WFI, so any
interrupt at or below TX_PORT_BASEPRI fails to wake the processor at all.
On a system where every interrupt is managed by ThreadX, that is every
interrupt, and the processor stays in WFI or in the low power mode entered
by tx_low_power_enter until something outside the mask happens.

Set PRIMASK and clear BASEPRI for the duration of the WFI, then restore
BASEPRI and clear PRIMASK. Interrupts remain masked across the whole
window, so the original race is still closed, but the wake-up condition is
now evaluated with BASEPRI clear and any enabled interrupt can end the wait.
A pending interrupt is taken at the CPSIE i, or at the existing unmask
further down the loop if it falls below TX_PORT_BASEPRI.

The change is made in the ARMv7-M and ARMv8-M sources under ports_arch,
including the module manager, and propagated to the generated ports with
the copy scripts. The ghs and keil ports do not implement
TX_PORT_USE_BASEPRI, and Cortex-M0/M0+/M23 have no BASEPRI register, so
those ports are unaffected.

Verified by assembling every patched GNU port for Cortex-M3, M4, M7, M33,
M55 and M85, with and without TX_PORT_USE_BASEPRI, TX_LOW_POWER and
TX_ENABLE_EXECUTION_CHANGE_NOTIFY, and by checking the disassembly of the
idle loop.

Fixes #279

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-09 11:29:14 -04:00
Frédéric Desbiens 146d57b235 Fixed the RISC-V64 trap frame size mismatch in the regression test BSP (#708)
The RISC-V64 port moved its interrupt frame to 528 bytes (65 slots plus 8
bytes of padding, so sp stays 16-byte aligned at a call) and published the
size as TX_RISCV_TRAP_FRAME_SIZE. The port sources were converted to use
it, but the shared regression test BSP was not: its trap_entry still
allocated a hardcoded 65 * REGBYTES, or 520 bytes.

Every interrupt therefore unwound 8 bytes more than it allocated:

    trap_entry:                   addi  sp,sp,-520
    _tx_thread_context_restore:   addi  sp,sp,528

On the RISC-V64 regression suite that left 24 of 95 tests failing in the
default configuration, typically as an illegal instruction once execution
reached a corrupted frame. The example BSP under the port directory was
converted with the port and was unaffected, which is why the functional
QEMU test kept passing.

The test BSP now takes both frame sizes from the port it is linked
against, so the two cannot drift apart again. A port that publishes no
contract keeps the historical layout, so the RISC-V32 side is unchanged
until its own port publishes one.

TX_RISCV_TRAP_CALL_FRAME_SIZE is restored to the RISC-V64 tx_port.h. It
was removed as unused when the frame sizes were introduced, but it is
part of the same contract: it is the space a trap entry reserves around a
call into C, and the psABI requires 16 bytes there rather than one
register slot.

Verified on QEMU with every linkable test built, comparing against the
commit before the port change:

    before the port change   2 failures out of 95 (both unlinkable)
    current dev              24 failures out of 95
    with this change          2 failures out of 95 (both unlinkable)

The two remaining failures predate all of this: newlib pulls _impure_ptr
out of R_RISCV_HI20 range for time(), so those two binaries do not link.
RISC-V32 is unchanged at 2 failures across all five configurations.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-09 10:56:57 -04:00
Frédéric Desbiens 9a03838381 Stopped the test runners from reporting an incomplete build as test failures (#709)
The cmake test runners drove Ninja with its default keep-going of 1, so the
first failing target ended the build. Every target scheduled after it was
simply absent, and ctest reports a missing binary as a failing test. A
single link error therefore produced a failure count that moved with build
scheduling order rather than with the code.

Measured on the RISC-V64 regression suite, where two targets genuinely
cannot link:

    before   79 of 95 test binaries built
    after    93 of 95 test binaries built

Fourteen perfectly good binaries were being skipped and counted as
failures. Passing -k 0 lets Ninja finish everything it can; the build still
exits non-zero when a target fails.

Three related problems in the same paths are fixed with it.

A failing configuration used to abort the loop over configurations, so
under set -e the ones after it went unbuilt or untested. The build loops
and the serial test loops now accumulate status and return it at the end,
which is what the parallel test branch already did with wait, and what
cmake_bootstrap.sh already documented for ctest.

Capturing that status removes the set -e protection inside the functions,
so two latent faults become reachable and are closed here. A failed pushd
would have let ctest run in the source tree, where it finds no tests and
reports success; the pushd is now guarded. And ctest's status was
discarded by the popd that follows it, so a configuration with failing
tests returned 0 and was reported as a pass; the status is now carried
past the popd and the summary steps.

Verified on the RISC-V32 suite, which has two genuine failures in each of
its five configurations. Both the serial and the parallel branch now test
all five and exit 8, where the serial branch previously stopped after the
first configuration.

The tx and smp runners are symlinks to scripts/cmake_bootstrap.sh, so they
are covered by the one change there.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-09 10:56:28 -04:00
Frédéric Desbiens 37f48fa83c Removed FIFO queueing from the ARMv7-A SMP ports so an ISR can no longer deadlock waiting for protection (#707)
The Cortex-A5, A7 and A9 SMP ports guarded the inter-core protection with
a FIFO wait list: a core that could not take the lock added itself to a
queue, incremented _tx_thread_smp_protect_wait_counts[core], and only the
core at the head of the queue was allowed to acquire.

Two parts of that scheme require the waiting core to be interruptible.
_tx_thread_smp_unprotect() refuses to release the protection while the
releasing core's own wait count is non-zero, and a queued core is taken
back out of the list by _tx_thread_context_restore() when it is preempted.
Neither can happen on a core that is spinning inside an ISR with
interrupts already masked, because the spin loop restores the caller's
interrupt posture rather than enabling interrupts. The core stays in the
list forever and the system deadlocks with cores stuck in the wait loop.

This is the deadlock Microsoft removed from the ARMv8-A SMP ports in
6.1.11, by dropping the wait list and using a plain LDAXR/STXR spinlock.
The same removal was announced for the ARMv7-A ports at the time but was
never made, so those three ports have carried the deadlock ever since.

The FIFO queueing is now removed from the ARMv7-A SMP ports as well.
_tx_thread_smp_protect() becomes an LDREX/STREX spinlock that releases
interrupts between attempts, matching the sequence already used by the
Cortex-R8 SMP port; _tx_thread_smp_unprotect() no longer consults the
wait counts; and _tx_thread_context_restore() no longer has to unqueue a
preempted core. The now unreferenced wait list macro headers are deleted.

Verified by assembling all three GNU ports with arm-none-eabi-gcc for
cortex-a5, cortex-a7 and cortex-a9, with and without
TX_ENABLE_FIQ_SUPPORT, TX_ENABLE_WFE and TX_MPCORE_DEBUG_ENABLE, and by
reading back the disassembly of the new protect sequence. The port
consistency checks pass.

Fixes #219

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-09 10:29:51 -04:00
Frédéric Desbiens d3fb72b6dd Aligned the simulator ports' fake stack pointer so ThreadX no longer performs misaligned ULONG accesses (#705)
The Linux, Win32 and Win64 simulation ports build a fake initial stack
pointer by subtracting a fixed 8 bytes from tx_thread_stack_end. That
field addresses the last byte of the thread's stack area, so it is one
less than an aligned address and the resulting pointer is misaligned by
construction, no matter how well aligned the stack the application
supplied was.

Two ULONG accesses then use that pointer. _tx_thread_stack_build() itself
clears the word below it, and _tx_thread_create() copies it into
tx_thread_stack_highest_ptr, which TX_THREAD_STACK_CHECK dereferences on
every suspend and resume when TX_ENABLE_STACK_CHECKING is defined. Both
are undefined behaviour. They happen to work on x86 but are reported by
GCC's undefined behaviour sanitizer, and would fault on a host that
requires natural alignment.

The fake stack pointer is now rounded down to a ULONG boundary, which
leaves it inside the stack area and makes both accesses aligned.

Verified by building the Linux port with -fsanitize=undefined and
running the demo: the two reported diagnostics are produced before the
change and neither appears after it. The tx and smp regression suites
pass, 98 and 114 tests respectively.

Fixes #218

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-09 09:34:33 -04:00
Frédéric Desbiens 2e48ce0d1b Removed the stray build artifacts committed with the RISC-V64 port fix (#706)
A local CMake build tree was staged by mistake alongside the RISC-V64
spec compliance work in #698, putting nine generated files on dev,
including a compiled kernel.elf, build.ninja, the Ninja dependency logs
and a QEMU run log. None of it belongs in the repository.

The root .gitignore listed build directories by name rather than by
pattern, so build/, build_qemu/, build_m7/ and the build_r52 variants
were covered but a differently named tree was not. Those entries are
replaced with a single build*/ pattern, which covers every existing name
and any future one. No tracked file matches the new pattern.

The artifacts remain reachable in history; only the working tree is
corrected, since rewriting a shared branch is the greater harm.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-09 09:31:00 -04:00
Frédéric Desbiens 515ab8aba1 Added the missing memory barrier so ARMv7-A SMP schedulers no longer miss preemptions (#704)
_tx_thread_schedule stores the newly selected thread into
_tx_thread_current_ptr[core] and then reloads _tx_thread_execute_ptr[core]
to detect a concurrent scheduling decision made by another core. On the
other side, _tx_thread_smp_core_interrupt stores the execute pointer and
then reads the current pointer to decide whether an inter-core interrupt is
required. This is a store-buffer pattern: without a barrier both cores can
observe stale values, the interrupt is skipped, and a ready thread with the
highest priority is never scheduled.

The barrier was added to the ARMv8-A SMP scheduler in 6.2.1, but the
ARMv7-A SMP ports were left untouched even though they implement the same
protocol. This adds the corresponding DMB to the Cortex-A5, Cortex-A7 and
Cortex-A9 SMP schedulers for both the AC5 and GNU toolchains.

The GNU variants were verified by assembling them with arm-none-eabi-gcc
13.2.1 for their respective cores.

Refs #209

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-09 09:10:36 -04:00
Frédéric Desbiens e99f0d4207 Corrected the module kernel stack size so it no longer overstates the usable stack (#701)
The module manager recorded tx_thread_module_kernel_stack_size as the raw
TXM_MODULE_KERNEL_STACK_SIZE constant, but the end of the kernel stack is aligned
downwards to an eight-byte boundary while _txm_module_manager_object_allocate only
guarantees ULONG alignment. The recorded size could therefore overstate the usable
stack by up to seven bytes. The scheduler copies this value into tx_thread_stack_size
whenever a user mode module thread enters the kernel, so the overstated value is
visible to RTOS-aware debuggers and to anything built on it.

The size is now derived from the aligned end minus the start. Also documented that
TX_ENABLE_STACK_CHECKING is not supported for module threads.

Refs #181

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-08 16:44:48 -04:00
Frédéric Desbiens 945f5f5caa Moved the ARC ISR enter callout onto the system stack (#700)
_tx_thread_context_save() calls _tx_execution_isr_enter when
TX_ENABLE_EXECUTION_CHANGE_NOTIFY is defined. On the path where an interrupt
preempted a running thread, that call was made before the switch to the system
stack, so the 32-byte call frame and the whole stack footprint of the callout
were taken from the interrupted thread's stack, on top of the 160-byte interrupt
frame the port had already allocated there.

The callout is supplied by the application, so its stack usage is not bounded by
ThreadX, and it is charged to every thread that happens to be running when an
interrupt arrives.

The two other callout sites in the same routine, the nested save and the idle
system save, already run on the system stack, as does the _tx_execution_isr_exit
call in _tx_thread_context_restore. _tx_thread_schedule was reordered in 6.1.9
so that _tx_execution_thread_enter runs on the system stack rather than the
thread stack; the same reorder was never applied to the context save.

The switch to the system stack now happens before the callout in the ARCv2_EM,
ARC_HS and SMP ARC_HS ports. _tx_thread_context_fast_save is unchanged because
the fast interrupt path never switches stacks by design.

Fixes #149

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-08 15:42:39 -04:00
Frédéric Desbiens b0ec8bfbb9 Fixed the invalid module data pointers in the absolute module load (#699)
An absolutely located module has its code and its data placed at two
independent fixed addresses by the module's linker script. The module
preamble carries the code and data sizes but not the data address, so
_txm_module_manager_absolute_load() could not determine where the
module's data area was. It computed txm_module_instance_data_start
from the code size and the preamble size, which yields a size rather
than an address, and it set txm_module_instance_module_data_base_address
one past the end of the byte pool allocation.

Added _txm_module_manager_absolute_load_extended(), which accepts the
module's data area address from the caller. Deprecated
_txm_module_manager_absolute_load(), which now forwards to the extended
service with an unknown data area location and rejects modules that
request memory protection, since the memory protection hardware cannot
be programmed to cover an unknown data area.

Fixes #450

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-08 10:03:45 -04:00
Frédéric Desbiens d7789f0b12 Fixed the clobbered return address in the RISC-V context save (#696)
_tx_thread_context_save() returns to its caller with ret, which uses the
return address held in ra. When TX_ENABLE_EXECUTION_CHANGE_NOTIFY was
defined, the call to _tx_execution_isr_enter overwrote ra with the address
of the instruction following the call, so the subsequent ret returned into
_tx_thread_context_save itself instead of the interrupt service routine.

The return address is now saved on the stack around the call and recovered
afterwards, which is the same idiom already used by the Arm ports. The fix
covers all three affected paths (nested save, thread save and idle system
save) in the risc-v32 GNU, risc-v32 IAR and risc-v64 GNU ports.

Fixes #348

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-03 14:17:49 -04:00
Frédéric Desbiens 508af549da Fixed the incorrect loop bound constant in the IAR file lock support (#695)
The IAR multithreaded library support code allocates its file lock
mutexes from an array of _MAX_FLOCK entries, but the wrap-around check
and the exhaustion check in __iar_file_Mtxinit() both compared against
_MAX_LOCK, the bound of the unrelated system lock mutex array.

When _MAX_FLOCK is greater than _MAX_LOCK, the free mutex index wrapped
early and the exhaustion check reported failure while free entries
remained, so *m was set to TX_NULL and the application faulted the first
time a file lock was taken.  When _MAX_FLOCK is smaller than _MAX_LOCK,
the free mutex index was allowed to run past the end of
__tx_iar_file_lock_mutexes and the exhaustion check could never fire.

Corrected all four comparisons in each of the 27 copies of tx_iar.c.

Fixes #444

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-03 13:50:31 -04:00
Frédéric Desbiens 17ff21a07f Fixed the garbled comment in the POSIX condition variable sources (#694)
The comment above the internal semaphore lookup in the POSIX condition
variable implementation was garbled: it contained a capitalisation typo
("COndition") and an incomplete sentence ("into a semaphore  a cast").
The four affected sources now share a single, correct wording.

This is a comment-only change; there is no functional impact.

Fixes #419

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-03 12:58:26 -04:00
Frédéric Desbiens c469a7756b Fixed the missing immediate prefix on MOV in Cortex-M schedulers (#693)
The BASEPRI-masking path in tx_thread_schedule wrote "MOV r0, 0" rather
than "MOV r0, #0". UAL requires the "#" prefix on an immediate operand.
GNU as and the LLVM-based assemblers accept the unprefixed form and emit
the intended encoding, but stricter assemblers reject it outright, so the
affected ports could not be built with those toolchains.

The GNU and AC6 sources had already been corrected; this brings the IAR
and AC5 sources into line. Verified that both spellings assemble to the
same Thumb-2 encoding (f04f 0000), so this is a source-correctness fix
with no change in generated code or runtime behaviour.

Covers 31 occurrences across the Cortex-M3, M4, M7, M33, M52, M55 and M85
ports, their module manager counterparts, and the shared ARMv7-M and
ARMv8-M architecture sources.

Fixes #461

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-03 11:44:27 -04:00
Frédéric Desbiens 13c8c768c7 Added the boot-at-EL1 option to the S32Z280 entry path (#690)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The Armv8-R AEM FVP entry path has carried TX_R52_BOOT_AT_EL1 since it
was written, for the case its own comment describes: "an earlier boot
stage or a vendor EL2 monitor has already dropped privilege to EL1".
This board's entry path did not, so a kernel could not be built as a
guest on it at all -- and it is the board where that matters most,
because it is the one with silicon behind it.

The bracket is the whole change.  Everything from the Thumb reset
trampoline to the ERET goes inside #ifndef TX_R52_BOOT_AT_EL1, and the
#else supplies a one-instruction A32 _start that branches to el1_entry.

A32 AND NOT T32, which is the one real difference from the standalone
entry.  The core resets in Thumb state here because the RTU boot
instruction NXP plants is a T32 branch, but a guest is not reached by
reset: it is reached by the monitor's ERET, and the monitor chooses the
state through SPSR.T.  Get the two out of agreement and the guest dies
on its first instruction with an undefined-instruction exception, which
looks exactly like a bad entry address and sends the reader to the
loader instead of to the ERET.

WHAT THE MONITOR INHERITS is enumerated at the #ifndef, next to the code
it replaces rather than in a document, because that is where somebody
adding a third board will be looking.  This board's EL2 block is
considerably larger than the model's, and each item on the list is
something a guest at EL1 provably cannot do rather than something it
merely does not: CNTFRQ is writable only at the highest implemented
exception level and reads zero out of reset; HCPTR.TCP10/TCP11 reset
set, trapping every EL1 floating-point access; HSCTLR.TE is an EL2
register (SCTLR.TE is EL1's, and el1_entry still clears it below);
ICC_HSRE.SRE makes every other ICC_* and ICH_* register exist at all;
the low-latency peripheral port enables reset to zero and an EL1 write
to that register traps to EL2; and the TCM enables are per-core with
ENABLEEL2 SILENTLY IGNORED from EL1 -- measured on both BTCM and CTCM,
the base took and bit 0 took while bit 1 stayed clear.

CNTHCTL.PL1PCTEN and PL1PCEN are the deliberate omission from that list,
and the note says why.  This path opens both, because a standalone
kernel owns the physical timer.  A monitor that TIME-partitions its
guests must not: a partition's physical time keeps running while it is
descheduled, so a guest reading it can observe that it was not running.
That is the monitor's decision rather than this file's, which is why the
list says what a guest cannot do rather than what a monitor should.

Verified both ways.  The three standalone images build and link
unchanged, and a kernel built with the option boots at EL1 on a
S32Z280-594EVB under an EL2 monitor, runs two threads through a queue
and a semaphore, and reports back -- with no other change to the kernel
or to its port.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-02 19:08:24 -04:00
Frédéric Desbiens 8c681c188e Asserted that the Cortex-R52 port refuses the options it documents as refused (#687)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
ports/cortex_r52/gnu/CMakeLists.txt rejects two option combinations at
configure time -- TX_R52_ENABLE_VFP without TX_R52_FLOAT_ABI=hard, and
TX_R52_ENABLE_FIQ_NESTING without TX_R52_ENABLE_FIQ.  Nothing exercised
either.  A guard that has stopped firing looks exactly like a guard nobody
has tripped, so both could have been silently disabled by a typo in a
variable name at any point and no check would have noticed.

check_gcc.sh gains a sixth stage that configures each rejected combination
and asserts the configure fails.  It needs no new harness: the script already
runs cmake as a subprocess for the CMake example builds, and the port's
CMakeLists is included by the toolchain file alone, so no example flags are
needed and the two negative configures stop almost immediately.

Two things the stage does that a thinner version would not.

It asserts the message text, not just the exit status.  A configure that
fails for an unrelated reason would otherwise be recorded as a guard doing
its job.

It also configures the SUPPORTED combination and requires that to succeed.
Two negative assertions on their own are satisfied by a guard that rejects
everything -- the port would be unbuildable and the check would still pass.
The positive case is what separates a guard that works from one that is
merely always on.

Verified negatively, four deliberate breaks, each caught:

  VFP guard condition forced false      -> "configure succeeded, but this
                                           combination cannot build"
  FIQ nesting guard forced false        -> the same, on that case
  guard fires, message text changed     -> "configure failed, but not on the
                                           expected guard"
  VFP guard condition forced true       -> "the supported combination was
                                           refused"

The fourth is the one the positive case exists for and the only one a
refusals-only stage would have missed.  ports/cortex_r52/gnu/CMakeLists.txt
was confirmed byte-identical to dev afterwards.

Full run passes: exit 0, all six stages, and --asm-only correctly skips the
new one.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-01 09:42:38 -04:00
Frédéric Desbiens 57390fc0fe Refused the Cortex-R52 VFP option without a hard float ABI (#686)
TX_R52_ENABLE_VFP with the default soft float ABI cannot build.  The option
defines TX_ENABLE_VFP_SUPPORT, which enables the VMRS, VSTMDB and VLDMIA
blocks in the port assembly, and -mfloat-abi=soft leaves the assembler with
no FPU to accept them.  The configuration fails with eight errors of the form

  tx_thread_system_return.S:118: Error: selected processor does not support
  `vmrs r4,FPSCR' in ARM mode

none of which mentions the float ABI, so a user has to reason from VMRS back
to the option that enabled it.

The guard that exists said otherwise.  It warned that "the compiler will not
emit floating-point instructions, so the VFP context path will never be
exercised", which describes a build that succeeds and is merely pointless --
and then let configure finish, so the warning scrolled past well before the
assembler errors appeared.

It is now a FATAL_ERROR that names the fix, which is what the same file
already does six lines above for TX_R52_ENABLE_FIQ_NESTING without
TX_R52_ENABLE_FIQ.  That combination is rejected for being "meaningless",
while this one, which cannot assemble at all, was only warned about.  The
severities were the wrong way round.  The option's definition also moves
below the check, so the block reads like the FIQ nesting one.

The ABI is not promoted to hard automatically.  TX_R52_FLOAT_ABI is a cache
variable the user may have set deliberately, and silently overriding an
explicit choice is worse than refusing a combination that cannot work.

readme_threadx.txt carried the same claim, and its option list marked the FIQ
nesting dependency inline but not this one.  Both corrected.

No regression test.  Nothing in the tree asserts a configure-time failure --
there is no harness for it, and the sibling FIQ nesting guard has none either
-- so a test for this would have to introduce that mechanism for one case.
Verified by hand in both directions instead: the soft-ABI combination now
stops at configure with the message above, and the hard-ABI feature build
(VFP, FIQ, IRQ nesting, FIQ nesting) builds its eight images clean and passes
ctest 8/8 on the Armv8-R AEM FVP.  scripts/check_gcc.sh passes unchanged; it
configures the Cortex-R52 CMake stage without TX_R52_ENABLE_VFP, so the new
branch is not on its path.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-01 08:48:20 -04:00
Frédéric Desbiens a2800fef16 Covered the SMP suspension teardown and the long byte pool search (#677)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
Against the merged SMP coverage report -- every build configuration instrumented
and unioned -- sixty-four lines of common_smp/src were uncovered, 5114 of 5178.
Fifty-three of them are closed here and the report reads 5167 of 5178. The SMP
coverage floor goes from 98 to 99 with it.

Thirty-six of the sixty-four were one loop repeated four times: the walk in
tx_block_pool_delete, tx_byte_pool_delete, tx_event_flags_delete and
tx_queue_delete that releases every thread suspended on the object with
TX_DELETED. The suite deletes all four object types after every single test, and
that is exactly why the loop never ran. test_control_cleanup in the ThreadX
suite deletes the application's objects first and its threads last, so a test
that ends with a thread parked on a queue has that thread walked out of it by
tx_queue_delete. The SMP suite's cleanup deletes the threads first, deliberately
-- it was changed so that no application-owned object is still referenced when
the object loops run, which is what stopped a class of teardown hang. The side
effect is that all four deletes now run against an empty suspension list.
tx_semaphore_delete is the one member of the family that was already covered,
because threadx_semaphore_delete_test deletes a busy semaphore on purpose.

threadx_object_delete_suspension_test is the same idea for the other four. Two
threads suspend on each of a block pool, a byte pool, an event flags group and a
queue; the control thread waits on the object's own suspended count through
tx_*_info_get rather than on an ordering it cannot guarantee across four cores,
deletes the object, and checks both waiters came out with TX_DELETED. Two
waiters rather than one so the loop takes its back edge as well as its body, and
every wait is bounded in ticks so a suspension that never arrives fails the test
instead of hanging it.

threadx_trace_entry_update_test and threadx_thread_misaligned_stack_test are
ports of the two tests that closed the equivalent gaps in common/src, and close
fourteen more lines here: tx_block_allocate 123, 175, 182, 319 and 326,
tx_byte_allocate 130, 210, 217, 359 and 366, tx_thread_system_suspend 504 and
560, tx_trace_object_register 221, and tx_thread_create 133. The one substantive
change is core confinement. The trace test needs thread 0 to suspend and thread
1 to then release what it waits for; on four cores thread 1 gives the block back
before thread 0 has suspended and the update block behind the suspension is
never reached, so both threads are excluded from cores 1 to 3. The misaligned
stack test needed no such change.

threadx_byte_memory_long_search_test closes three of the eleven in
tx_byte_pool_search. Lines 264, 267 and 270 are the
TX_BYTE_POOL_MULTIPLE_BLOCK_SEARCH limit -- twenty on this port -- where a long
search drops and retakes protection so that it cannot lock the other cores out
for the whole walk. No byte pool in the suite ever had twenty fragments. This
one is filled with small chunks until it refuses another and then has every
second chunk released, so the free fragments are never adjacent and cannot be
merged, and the request is larger than any of them but smaller than the pool's
theoretical total, which is what makes _tx_byte_pool_search walk rather than
refuse at the door. The layout is asserted rather than assumed: the test checks
the fragment count and checks the probe request really does fail before the
workers start, because either would otherwise turn it into a silent no-op.

Eleven lines remain and they are not a to-do list. Eight are the delay loop in
tx_byte_pool_search that fires when another thread claims the pool inside the
window the search opens. The Linux SMP port serialises all four cores on one
pthread mutex, so that window is an unlock immediately followed by a lock on
that mutex, and glibc hands an uncontended mutex straight back to the thread
that just released it: measured over 180,003 windows across three cores, with
zero handovers. The shipped test therefore does 250 searches per worker rather
than the sixty thousand that probe used, because the twenty-block threshold is
crossed by the first search. The other three are in tx_thread_smp_utilities.
Line 149 is a range guard placed after the shift it is meant to guard, so
reaching it needs a shift by the width of the type; the fix is to move the check
above the shift, matching the TX_MAX_PRIORITIES > 32 variant of the same
function, and that belongs in its own change. Lines 1073 and 1074 need a mutex
owner that is genuinely executing on another core when a waiter suspends, and
three shapes were tried without producing one on this port.

Measured twice before and twice after, every gcda deleted between runs and
570 of 570 tests passing each time: 5114 of 5178 both times before, 5167 of 5178
both times after. Branch coverage goes from 2768 to 2821 and 2823 of 3548. A
floor of 99 needs 5127, so the ratchet lands with forty lines of headroom
against a numerator that has been seen moving by two between runs.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-28 10:41:39 -04:00
Frédéric Desbiens ff7fbe02f8 Added a GCC check for the Arm ports and ran it in CI (#675)
* Assembled the module ports, which no check had ever compiled

scripts/check_clang.sh globbed ports_module/*/gnu/src, which does not exist --
the module ports keep their assembly in module_manager/src. The [ -d ] guard
skipped it in silence, so 116 assembly files across nine Arm module ports were
assembled by no check, with either compiler, in the script whose own comments
state three times that "a port that is simply absent from the count reads as
covered". Stage 1 goes from 724 of 724 to 840 of 840; the feature-macro stage
had the same gap and goes from 412 files to 469.

Correcting the path exposed five defects, and only one of them was a build
failure. The other four assembled cleanly and did the wrong thing, because GAS
runs the C preprocessor on .S and not on .s:

  ports_smp/cortex_a7_smp/gnu/src/tx_thread_smp_unprotect.s, the only .s in a
  directory of twenty-one .S, ignored all four of its own feature macros. It
  wrote the caller's LR into the protection structure on every unprotect -- a
  store guarded by TX_MPCORE_DEBUG_ENABLE -- sent an unconditional SEV, and
  returned through both BX lr and MOV pc, lr. Its cortex_a5_smp and
  cortex_a9_smp siblings are .S.

  ports_module/cortex_m33/.../tx_thread_stack_build.s emitted both arms of an
  #ifdef TX_SINGLE_MODE_SECURE, so the non-secure LR value overwrote the secure
  one and the secure build got the wrong frame.

  ports_module/cortex_m23/.../tx_thread_context_{save,restore}.S carried the
  POP {r0, lr} that check_clang.sh's own comment describes as the reason the
  feature-macro stage exists. The 16-bit Thumb POP takes r0-r7 and pc only.
  The identical fix already sits in ports/cortex_m23/gnu/src; the module copy
  never got it because nothing scanned it.

  ports_module/cortex_m23/.../tx_thread_secure_stack_initialize.S used MOV
  rather than MOVS for an 8-bit immediate, latent behind TX_SINGLE_MODE_SECURE.
  Both siblings in the same directory already use MOVS.

  ports_module/cortex_a7/gnu/module_manager/src is the one that failed to
  assemble, on GCC 14.3 as well as on LLVM: #define SYS_MODE was never
  expanded, so #SYS_MODE reached the assembler as an undefined symbol.

Twenty-nine .s files under gnu trees are renamed to .S. Every one of them is
already named .S by the build scripts that compile it, so this repairs those
scripts rather than churning them -- ports_module/cortex_a7's build_threadx.bat
names all eighteen with a capital S, and works today only on a case-insensitive
filesystem. Renaming rather than converting the #defines to GNU assignments is
what fixes the #ifdef blocks as well as the constants; the assignments would
have fixed two files and left twenty-seven silently ignoring their macros.

Files with no preprocessor directives are left as .s: they are not broken, and
check_ports.sh gains a check that keeps them that way. Only the gnu trees are
checked there -- the IAR, Arm Compiler 5 and Keil assemblers preprocess .s
themselves, and about three hundred files in this repository rely on that.

Verified with both toolchains on the same tree: 840 of 840 assembled by
ATfE 22.1.0 and by arm-gnu-toolchain 14.3.rel1, all five stages of
check_clang.sh green, and check_ports.sh green including the reproducibility
check. The new check was shown to fail by planting a copy of the file it was
written for.

No regression test accompanies this. The assembly it covers is executed by no
host test, and the check itself going from 724 files to 840 is the coverage
AGENTS.md asks for -- together with the new check_ports.sh section, which is
what stops the class recurring.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Fixed the AArch64 samples, none of which had ever linked with GCC

Every AArch64 gnu example build failed at the sample link, all 27 of them --
13 under ports/ and 14 under ports_smp/:

  libg.a(libc_a-init.o): in function `__libc_init_array':
      undefined reference to `_init'
      relocation truncated to fit: R_AARCH64_CALL26 against undefined
          symbol `_init'
  libg.a(libc_a-fini.o): in function `__libc_fini_array':
      undefined reference to `_fini'

build_threadx_sample.sh links with -nostartfiles, which is correct for a port
carrying its own reset path, and that drops crti.o and crtn.o along with
everything else. startup.S calls __libc_init_array by design, and newlib's
implementation calls _init, which crti.o is what defines. The AArch32 scripts
are unaffected: they use nosys.specs and never reach __libc_init_array.

The fix links crti.o and crtn.o explicitly, bracketing the object list -- the
first must precede every .init contribution and the second must follow all of
them, so their position is load-bearing rather than stylistic. Both paths come
from the compiler's own -print-file-name, so nothing here hard-codes a
toolchain layout.

The atfe branch sets both to empty, deliberately: picolibc's __libc_init_array
does not call _init, those 27 images link today, and adding crti.o would change
a working link for no reason. That is also why check_clang.sh is green on these
and does not list them as expected to fail -- the LLVM path never reached the
gap, so nothing has ever linked them and failed.

Fixed in ports_arch/ARMv8-A/threadx/ports/gnu/example_build, which is the
single source for both the ports/ and ports_smp/ copies, then regenerated with
update.sh --port-sets tx,tx_smp. The 27 generated copies are in this commit
because ports_arch_check compares them.

Verified: all 27 link with arm-gnu-toolchain 14.3.rel1 aarch64-none-elf, where
0 of 27 did before; _init and _fini disassemble to the expected crti prologue
and crtn epilogue over a ret; check_clang.sh with ATfE 22.1.0 is still green on
all five stages, including the 42 script-driven example builds; check_ports.sh
is green including the reproducibility check.

No regression test: these are link-only example images that no host test
executes. What guards them is check_clang.sh's example stage today, and
check_gcc.sh's, which is the next change and is the reason this was found.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Added a GCC check for the Arm ports, which nothing had ever compiled

GCC is the project's declared default compiler (AGENTS.md, "The default
compiler for the project is GCC 14 on Linux"), it is what the gnu ports exist
for, and it is what nearly every downstream user builds with -- and nothing in
CI compiled a line of any port with it. The only cross-compilation check that
ran was the LLVM one, so the ATfE path was better guarded than the GNU one, on
ports whose directory is literally named gnu. ci_cortex_m covers four port
families; this covers forty.

Five stages, mirroring scripts/check_clang.sh stage for stage:

  1. assemble every .S and .s of every Arm gnu port -- 840 files
  2. assemble again the parts behind TX_ENABLE_VFP_SUPPORT,
     TX_ENABLE_FIQ_SUPPORT, TX_LOW_POWER and
     TX_ENABLE_EXECUTION_CHANGE_NOTIFY -- 469 files
  3. compile common/src for one core per architecture profile -- 185 x 9
  4. link the script-driven example builds -- 42
  5. link the CMake-driven Cortex-R52 images -- 5

Two scripts rather than one with a --toolchain flag: the flag surface differs
(a prefixed driver against --target=), the C library differs, and the set of
examples that can link differs. Folding them together makes it easy to weaken
one check while working on the other.

Two toolchains, both required. Arm ships arm-none-eabi and aarch64-none-elf as
separate downloads and PORT_TARGET maps every port to one of exactly those two
triples, so --arm-none-eabi and --aarch64-none-elf each take a driver or the
directory holding it, defaulting to the environment and then to PATH. A missing
one is a hard error rather than a soft skip: letting a run cover half the tree
and still report "all checks passed" is the failure this script exists to end.

PORT_TARGET is copied verbatim from check_clang.sh, including its warning not
to prefix-match core names -- cortex_a5* also matches the AArch64 cortex_a53.

VFP_EXTRA is the one map that is not a copy, and check_clang.sh's comment about
it is false for GCC. That comment says the A-profile defaults are already
correct; arm-none-eabi-gcc defaults to -mfloat-abi=soft, which disables the FPU
outright, so every VFP file fails with "selected processor does not support
'vmrs r1,FPSCR' in ARM mode". -mfloat-abi=hard alone is the fix and is the
right one, because it selects the core's own default FPU rather than naming a
-d16 one -- which is the trap the clang script warns about, since the
A-profile paths save D16-D31. Cortex-R4 is the exception in both scripts and
for the same reason: its FPU is an option rather than part of the core, so an
explicit -mfpu is required. Every value was measured against 14.3.rel1.

Stage 4 *unsets* TOOLCHAIN rather than setting it. The example build scripts
already default to GNU, and a stray TOOLCHAIN=atfe from a developer's shell
would otherwise make this stage silently check the other compiler. It cleans
the example directories on both sides, because the success test is the
existence of sample_threadx.out rather than the driver's exit status, and a
stale image from a previous toolchain would report success. Failure logs are
printed unfiltered: a missing tool says "command not found", and GNU ld's
undefined-symbol lines carry no "error:" at all.

Every skip is printed by name with a reason, per the house rule check_clang.sh
states three times -- a port simply absent from the count reads as covered.
This script also says outright that arm9 and arm11 are Arm and are skipped for
having no PORT_TARGET entry, which the clang script's "not Arm" wording glosses.

Verified on this tree with arm-gnu-toolchain 14.3.rel1: all five stages green,
every count identical to check_clang.sh's on the same tree -- 840, 469, 185x9,
42, 5 -- in 4m28s.

The failure paths were tested, not assumed. A deliberately broken .S in a
module port is reported by name and line in stages 1 and 2 and exits 1, in
--quiet mode as well. Reverting the AArch64 _init/_fini fix on one port only
gives "FAIL: cortex_a53: example build produced no image", 41 of 42, and exit
1 -- and the log tail it prints contains no "error:" anywhere, which is why it
is not filtered. A missing or wrong toolchain path exits 1 naming which triple
was not found.

RISC-V is deliberately out of scope for this first version: both ports
assemble 8 of 8 with the project's own cmake flags, but adding them widens the
toolchain download and the review surface for a family that is not regressing.

No regression test accompanies this. The script is the test, it exercises no
runtime behaviour, and its own failure paths are exercised above.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Ran the GCC port check in CI, on dev as well as master

scripts/check_gcc.sh with nothing invoking it would be a script nobody runs.
This adds the workflow, modelled on clang_check.yml, and fixes a trigger gap in
that file at the same time.

One job, two cache steps. Arm ships AArch32 and AArch64 as separate downloads
and the script needs both, so two caches keep the checks list short and let a
single invocation see both compilers. The AArch32 cache path and key match
cortex_m's exactly, so the two workflows share one entry rather than each
holding its own copy of the same archive -- noted in a comment, because the
only symptom of breaking that is a slower run.

Both triggers name dev. A workflow that triggers only on master gates no pull
request anybody opens; that is the defect ports_arch_check.yml carries a
comment about, and it cost cortex_m three months of failing in seven seconds
unnoticed. push is included as well as pull_request so dev's own history has a
baseline and a bad squash-merge is caught rather than waiting for the next PR.

The checksum suffix is .sha256asc and it is not interchangeable with .sha256.
Arm publishes both for this release, and verified 26 Aug 2026, the .sha256 file
for arm-none-eabi contains a 32-character MD5 rather than a SHA-256, so
sha256sum -c on it fails with "no properly formatted checksum lines found".
.sha256asc is a plain sha256sum-format line for both triples. The plan warned
that this suffix had changed between releases; the sharper truth is that both
suffixes exist simultaneously and one of them is not a SHA-256 at all. Recorded
in a comment beside the step.

Verified before writing them in rather than copied: both archive URLs and both
checksum URLs resolve, the archives are xz, the checksum files are
sha256sum-format for .sha256asc, and the AArch64 archive extracts to
arm-gnu-toolchain-14.3.rel1-x86_64-aarch64-none-elf/bin/aarch64-none-elf-gcc,
which is the path the workflow builds.

The paths: lists are duplicated between push and pull_request rather than
shared through a YAML anchor, deliberately: GitHub Actions' parser does not
dependably honour anchors and the failure mode is the workflow refusing to
parse, which is the cortex_m failure again. Ten duplicated lines are cheaper.

clang_check.yml's paths: list was missing CMakeLists.txt, cmake/ and
common_smp/, so that check did not run when files it reads changed -- the
ports_smp example builds compile common_smp/src and its CMake stage reads the
toolchain file and the top-level project. Both lists are now identical apart
from each file's own name, and both say so.

cortex_m is kept rather than deleted, against the plan's recommendation. It
builds four ports *through CMake*, and that is the only thing exercising
cmake/cortex_m*.cmake and the top-level CMakeLists for the M profile; this
script's CMake stage covers cortex_r52 only. The overlap is the assembly and
the C sources, not the build system, so deleting it would lose coverage rather
than remove a duplicate. Said so in the workflow header.

The script is passed explicit toolchain paths rather than left to find the
drivers on PATH, so nothing about the runner image can decide which compiler
runs, and it prints both versions it resolved.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-28 10:40:37 -04:00
Frédéric Desbiens 5b94bad6a2 Fixed the AArch64 samples, none of which had ever linked with GCC (#673)
Every AArch64 gnu example build failed at the sample link, all 27 of them --
13 under ports/ and 14 under ports_smp/:

  libg.a(libc_a-init.o): in function `__libc_init_array':
      undefined reference to `_init'
      relocation truncated to fit: R_AARCH64_CALL26 against undefined
          symbol `_init'
  libg.a(libc_a-fini.o): in function `__libc_fini_array':
      undefined reference to `_fini'

build_threadx_sample.sh links with -nostartfiles, which is correct for a port
carrying its own reset path, and that drops crti.o and crtn.o along with
everything else. startup.S calls __libc_init_array by design, and newlib's
implementation calls _init, which crti.o is what defines. The AArch32 scripts
are unaffected: they use nosys.specs and never reach __libc_init_array.

The fix links crti.o and crtn.o explicitly, bracketing the object list -- the
first must precede every .init contribution and the second must follow all of
them, so their position is load-bearing rather than stylistic. Both paths come
from the compiler's own -print-file-name, so nothing here hard-codes a
toolchain layout.

The atfe branch sets both to empty, deliberately: picolibc's __libc_init_array
does not call _init, those 27 images link today, and adding crti.o would change
a working link for no reason. That is also why check_clang.sh is green on these
and does not list them as expected to fail -- the LLVM path never reached the
gap, so nothing has ever linked them and failed.

Fixed in ports_arch/ARMv8-A/threadx/ports/gnu/example_build, which is the
single source for both the ports/ and ports_smp/ copies, then regenerated with
update.sh --port-sets tx,tx_smp. The 27 generated copies are in this commit
because ports_arch_check compares them.

Verified: all 27 link with arm-gnu-toolchain 14.3.rel1 aarch64-none-elf, where
0 of 27 did before; _init and _fini disassemble to the expected crti prologue
and crtn epilogue over a ret; check_clang.sh with ATfE 22.1.0 is still green on
all five stages, including the 42 script-driven example builds; check_ports.sh
is green including the reproducibility check.

No regression test: these are link-only example images that no host test
executes. What guards them is check_clang.sh's example stage today, and
check_gcc.sh's, which is the next change and is the reason this was found.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-28 09:26:12 -04:00
Frédéric Desbiens 9c32abb17d Assembled the module ports, which no check had ever compiled (#672)
scripts/check_clang.sh globbed ports_module/*/gnu/src, which does not exist --
the module ports keep their assembly in module_manager/src. The [ -d ] guard
skipped it in silence, so 116 assembly files across nine Arm module ports were
assembled by no check, with either compiler, in the script whose own comments
state three times that "a port that is simply absent from the count reads as
covered". Stage 1 goes from 724 of 724 to 840 of 840; the feature-macro stage
had the same gap and goes from 412 files to 469.

Correcting the path exposed five defects, and only one of them was a build
failure. The other four assembled cleanly and did the wrong thing, because GAS
runs the C preprocessor on .S and not on .s:

  ports_smp/cortex_a7_smp/gnu/src/tx_thread_smp_unprotect.s, the only .s in a
  directory of twenty-one .S, ignored all four of its own feature macros. It
  wrote the caller's LR into the protection structure on every unprotect -- a
  store guarded by TX_MPCORE_DEBUG_ENABLE -- sent an unconditional SEV, and
  returned through both BX lr and MOV pc, lr. Its cortex_a5_smp and
  cortex_a9_smp siblings are .S.

  ports_module/cortex_m33/.../tx_thread_stack_build.s emitted both arms of an
  #ifdef TX_SINGLE_MODE_SECURE, so the non-secure LR value overwrote the secure
  one and the secure build got the wrong frame.

  ports_module/cortex_m23/.../tx_thread_context_{save,restore}.S carried the
  POP {r0, lr} that check_clang.sh's own comment describes as the reason the
  feature-macro stage exists. The 16-bit Thumb POP takes r0-r7 and pc only.
  The identical fix already sits in ports/cortex_m23/gnu/src; the module copy
  never got it because nothing scanned it.

  ports_module/cortex_m23/.../tx_thread_secure_stack_initialize.S used MOV
  rather than MOVS for an 8-bit immediate, latent behind TX_SINGLE_MODE_SECURE.
  Both siblings in the same directory already use MOVS.

  ports_module/cortex_a7/gnu/module_manager/src is the one that failed to
  assemble, on GCC 14.3 as well as on LLVM: #define SYS_MODE was never
  expanded, so #SYS_MODE reached the assembler as an undefined symbol.

Twenty-nine .s files under gnu trees are renamed to .S. Every one of them is
already named .S by the build scripts that compile it, so this repairs those
scripts rather than churning them -- ports_module/cortex_a7's build_threadx.bat
names all eighteen with a capital S, and works today only on a case-insensitive
filesystem. Renaming rather than converting the #defines to GNU assignments is
what fixes the #ifdef blocks as well as the constants; the assignments would
have fixed two files and left twenty-seven silently ignoring their macros.

Files with no preprocessor directives are left as .s: they are not broken, and
check_ports.sh gains a check that keeps them that way. Only the gnu trees are
checked there -- the IAR, Arm Compiler 5 and Keil assemblers preprocess .s
themselves, and about three hundred files in this repository rely on that.

Verified with both toolchains on the same tree: 840 of 840 assembled by
ATfE 22.1.0 and by arm-gnu-toolchain 14.3.rel1, all five stages of
check_clang.sh green, and check_ports.sh green including the reproducibility
check. The new check was shown to fail by planting a copy of the file it was
written for.

No regression test accompanies this. The assembly it covers is executed by no
host test, and the check itself going from 724 files to 840 is the coverage
AGENTS.md asks for -- together with the new check_ports.sh section, which is
what stops the class recurring.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-28 09:24:30 -04:00
Frédéric Desbiens 147754cc86 Enforced a coverage floor on the merged report (#667)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The coverage summary reported a percentage and could not fail. Coverage
could fall from 99.97% to anything at all and every check stayed green,
against an AGENTS.md that asks for 100% test coverage -- a stated
requirement measured with a gauge that had no failure mode.

CodeCoverageSummary already takes thresholds and fail_below_min; neither
was set. Both are now, through a new coverage_thresholds input on the
template, because the two suites do not sit at the same figure: ThreadX
99, SMP 98.

Three things were probed against the pinned action on a runner before
picking those numbers, using the real merged.xml files from the dev push
run of #666.

The floor compares the line rate and nothing else. That mattered because
branch coverage is around 78% in both suites while line coverage is
98.8-100%, so a floor aimed at the line figure would have been an
immediate red wall had it tested branches or the lower of the two. The
ThreadX report at 100.00% lines and 77.67% branches clears a floor of 99.

The thresholds are whole numbers. '99.9 100' -- the value this was meant
to be -- is rejected with 'System.ArgumentException - Threshold parameter
set incorrectly.', and the step fails whether or not fail_below_min is
set. So the choice is 99 or 100 with nothing between.

100 would fail on a race. tx_thread_system_resume.c:529 is reached by
timing rather than by construction and flaps between runs of the same
green tree, which is why #666 left it; 4502/4503 fails a floor of 100 and
clears one of 99. A coverage gate that goes red on a coin toss is how
coverage gates get switched off.

SMP is 5114/5178 lines, 98.76%, with 64 uncovered lines across 11 files
of common_smp/src -- #666 closed the equivalent gaps in common/src only.
A shared floor of 99 would have failed that job on every run while
ThreadX passed.

One limit is recorded in the file rather than fixed: an empty report
reads as 100%. gcovr writes line-rate="1.0" beside lines-valid="0" when
it finds no data, and the action prints 'Line Rate = 100% (0 / 0)' and
passes any floor. The check for that is the emptiness assertion #664 put
in each suite's coverage.sh, not this one.

Also corrected two stale filenames in the deploy job's comment: since
#665 each coverage artifact carries merged.xml, not
default_build_coverage.xml. Verified on the runner.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 09:13:35 -04:00
Frédéric Desbiens f89d65f041 Covered the trace entry update paths and the misaligned stack adjustment (#666)
Against the merged coverage report -- every build configuration instrumented and
unioned -- eighteen lines of common/src were uncovered. Seventeen of them were
one missing scenario rather than eighteen separate gaps.

tx_block_allocate, tx_byte_allocate, tx_thread_system_suspend and
tx_thread_system_resume each carry blocks under TX_ENABLE_EVENT_TRACE that go
back and patch a trace entry once the call has done its work, all of the shape:

    if (entry_ptr != TX_NULL)
    {
        if (time_stamp == entry_ptr -> tx_trace_buffer_entry_time_stamp)

entry_ptr comes from _tx_trace_buffer_current_ptr, which stays TX_NULL until
tx_trace_enable is called at run time. Building with TX_ENABLE_EVENT_TRACE is
not enough, and exactly one test in the suite enables tracing --
threadx_trace_basic_test -- which tests the enable API itself and never calls
either allocator. So those blocks sat in the report's denominator and never in
its covered set.

threadx_trace_entry_update_test enables tracing and then drives both allocators
twice each, once on the path that succeeds immediately and once through a
suspension that a second thread satisfies, since each allocator carries one
update block on either side. It then sleeps so that the last runnable thread
suspends with nothing ready to take over: tx_thread_system_suspend lines 345 and
351 are on the branch that sets _tx_thread_execute_ptr to TX_NULL, and the two
allocator suspensions never reach it because the other thread was always ready.

The same test closes tx_trace_object_register's TX_NULL name break by creating a
semaphore with no name. A semaphore and not a thread deliberately: for
TX_TRACE_OBJECT_TYPE_THREAD the register function dereferences the pointer it is
given to read the thread's priority, so that type needs a real TX_THREAD behind
it. threadx_trace_basic_test makes the equivalent call only under
ifndef TX_ENABLE_EVENT_TRACE, against the no-op stub.

threadx_thread_misaligned_stack_test covers the remaining line,
tx_thread_create.c:136, where a stack that does not begin on a ULONG boundary
costs a ULONG of size so that rounding the start up cannot run past the end of
the caller's buffer. Every other test hands tx_thread_create an aligned stack.
That line is compiled only under TX_ENABLE_STACK_CHECKING, so it is absent from
three of the five configurations' reports rather than uncovered in them, and it
was verified under stack_checking_build.

Measured on the merged report, all 480 tests passing and 5 of 5 configurations
green: 4485 of 4503 covered before, 4502 of 4503 after, denominator unchanged.

One line remains, tx_thread_system_resume.c:529, and it is the report's last
flapping line rather than a standing gap -- two clean runs of the same tree gave
4503 of 4503 and 4502 of 4503. Reaching it by construction was tried twice and
failed both times, so it is left alone here. It needs the preempt disable flag
and the system state both clear, and tx_thread_resume raises the preempt disable
flag before calling _tx_thread_system_resume, as do the put and send paths;
creating a higher priority auto-start thread from thread context does not raise
it but does not reach the check either, which a probe showed is executed only
during initialisation, with the system state at TX_INITIALIZE_IN_PROGRESS.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 08:14:38 -04:00
Frédéric Desbiens 3e85bbd431 Instrumented every build configuration and merged their coverage (#665)
Only default_build_coverage carried -fprofile-arcs, because the gate was the
build type and it is the only one of five whose name ends in _coverage. The
other four build and run all their tests and their coverage was discarded. That
is not redundancy thrown away: each configuration selects a different set of TX_
feature macros, so the code the other four compile is absent from the
denominator rather than uncovered in it.

TX_COVERAGE instruments a build regardless of its name, defaulting to OFF so a
single configuration built by hand behaves as before. coverage.sh gains a
--merge mode that unions the per-configuration JSON tracefiles, and
cmake_bootstrap.sh runs it after the test loop so a local run produces the same
merged report CI reads. The template sets TX_COVERAGE for build and test, and
coverage_name moves to the merged report.

Measured on the ThreadX suite, all 480 tests passing:

  default_build_coverage           3827 valid   3827 covered
  disable_notify_callbacks_build   3767         3766
  stack_checking_build             3857         3856
  stack_checking_rand_fill_build   3862         3861
  trace_build                      4123         4108
  merged                           4503         4487

The denominator grows by 676 lines, 17.7%, and the figure moves from 99.97% to
99.64%. The second one is honest, and the drop is the point rather than a
regression: the denominator now includes code the old report never counted. The
union also contains a file the old report did not contain at all --
tx_thread_stack_error_handler.c compiles only under TX_ENABLE_STACK_CHECKING, so
it was not listed at 0%, it was simply absent. 177 files becomes 178.

Coverage collection moved out of test() and now runs after the test loop, one
configuration at a time. gcov writes its intermediate gcov files into the
directory gcovr is rooted at, and coverage.sh roots every configuration at the
repository root so filenames come out repo-relative. Five concurrent gcovr
processes therefore share one scratch directory and delete each other's output:
the first full run of this change passed all 480 tests and produced no report
for three of the five configurations. Measured both ways -- two gcovr rooted at
the repository root fail concurrently and succeed in sequence. CI would not
have caught it, because test_tx.sh sets CTEST_PARALLEL_LEVEL=1 and takes the
serial branch.

Per-configuration output moved under coverage_report/per_configuration/ and is
excluded from the Pages artifact. The deploy job merges the ThreadX and SMP
artifacts into one tree and every configuration directory has the same name in
both, so left at the top level one suite's would overwrite the other's on the
published site.

On the SMP suite, an earlier run of this change saw trace_build fail
threadx_smp_time_slice_test and then hang, which raised the question of whether
-fprofile-arcs perturbs a timing-sensitive test. It does not. Sixteen runs
settle it, and the decisive one is that threadx_smp_time_slice_test failed
ERROR #31 -- twice in a row under --repeat until-pass:2 -- on an uninstrumented
build, in the exact shape CI runs, while three instrumented runs of that shape
passed 5 of 5. In the CI shape, CTEST_PARALLEL_LEVEL=1 run.sh test all:

  TX_COVERAGE=OFF   3 runs   2 green, one ERROR #31        310 s
  TX_COVERAGE=ON    3 runs   3 green, 5/5 each             325-329 s

So the test is a pre-existing flake on dev and instrumenting all five costs
about 5% of the suite's wall clock. Separately, and also in both instrumented
and uninstrumented builds, run.sh's parallel branch -- what a developer gets
typing run.sh test all with no CTEST_PARALLEL_LEVEL -- hangs under its own load,
four times in twelve runs. Several SMP tests create 1024 ThreadX threads by
construction and the Linux port backs each with a pthread, so five
configurations at once put on the order of 5000 threads on the machine. CI sets
CTEST_PARALLEL_LEVEL=1 and does not take that branch.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 08:13:33 -04:00
Frédéric Desbiens b756220c43 Fixed the coverage report's paths and scoping (#664)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The Cobertura XML embedded absolute machine paths, and the flag that looked
like it scoped the report to one build configuration was doing nothing at all.
Both coverage.sh scripts had the defect; both are fixed here, because the SMP
report is published to the same Pages site as the ThreadX one.

Paths. -r was the build directory and -f pointed outside it, so gcovr could not
express the sources relative to the root and fell back to absolute paths. The
result named files as /home/runner/work/threadx/threadx/common/src/... while
the <source> element beside them said build/default_build_coverage, so the two
halves of the same file disagreed and nothing could map coverage back to the
repository. -r is the repository root now and -f an absolute path beneath it,
which gives filename="common/src/tx_block_allocate.c".

Both must be absolute: -r ../../.. -f common/src produces a report of zero
files and exits 0, which is the worst failure mode available here.

Scoping. --object-directory does not restrict which gcda files are found -- it
tells gcovr how to get from a gcda file back to the compiler's working
directory. Pointed at an empty directory it still produced the full 177-file
report. That was harmless only by accident, because -r build/$1 constrained the
search instead; moving -r to the repository root removes that accident, so the
two changes have to land together. Measured, with a second instrumented
configuration deliberately made sparser than the first:

  scoped by the positional search path   3221 of 3827 lines -- the truth
  no search path, -r at the repo root    3827 of 3827 -- silently merged
  --object-directory at the sparse tree  3827 of 3827 -- scopes nothing

So the search path is load-bearing, and it matters ahead of instrumenting all
five configurations: without it each configuration would have reported the
union as its own.

An empty report is not an error to gcovr -- it warns and exits 0 -- and it
carries line-rate="1.0" next to lines-valid="0", so a consumer reads no data at
all as fully covered. No coverage threshold can catch that, since an empty
report passes any threshold. Hence the explicit assertion that the report has
content, which fires with exit 1 on an object directory that exists but is
empty, where the old shape returned 177 files and exit 0.

Also says out loud that ports/linux/gnu/src is deliberately outside the filter.
gcno files exist for it and it is dropped without a word today.

Number-neutral, and that was the test. Over the same frozen gcda, changing only
the gcovr invocation: ThreadX 3827 of 3827 lines and 1993 of 1994 branches
across 177 files, SMP 4739 of 4791 and 2417 of 2430 across 185, before and
after alike. Same answers on gcovr 7.0, 8.3 and 8.6, so the change is not
wedged to the current pin. End to end through run.sh, 96 of 96 and 110 of 110
pass with the reports written.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 17:16:51 -04:00
Frédéric Desbiens 2d9b9f7417 Bumped gcovr off the 4.1 pin it had been held on since 2018 (#663)
The coverage tooling was pinned to gcovr 4.1, released in 2018, and that
version is missing the two options the coverage work needs next: --json and
--add-tracefile, which is how the five build configurations get merged into
one report. This moves the pin to 8.6, the current release. The pin stays
exact, and it stays hand-moved: it lives in a shell script, and no Dependabot
ecosystem can parse that.

Isolated deliberately, so that a movement in the coverage number caused by the
tool could not be confused with one caused by a later change. Measured on the
default_build_coverage tree of test/tx, over the same gcda with the same gcov,
varying only the gcovr version:

  gcovr    lines-valid  branches-valid  files
  4.1      3827         1994            177
  7.0      3827         1994            177
  8.3      3827         1994            177
  8.6      3827         1994            177

So the denominator does not move with the tool at all, and this bump moves no
number. The plan this came from expected 3822 to become 3827; that figure does
not reproduce, under gcc-13 or gcc-14, with or without --object-directory. The
only variant that changes the count is dropping the -f filter, which collapses
the report to nothing.

Two things found while measuring, both recorded because they matter to what
comes next.

The coverage numerator is not deterministic. On an identical tree with an
identical compiler, three consecutive runs of the full suite -- all 96 tests
passing every time -- reported 3826, 3827 and 3827 covered lines. The line that
flickers is tx_thread_system_resume.c:529, the preemption path of
_tx_thread_system_resume, and it takes its guarding branch with it. It has been
described as never executed; it is executed on some runs and not others. A
coverage floor has to be set with that in mind, and the honest fix is a test
that takes the path deliberately.

Reading gcc-13 output, the compiler the runners actually use, gcovr 8.6 runs
the existing coverage.sh unchanged: Cobertura XML and 181 HTML files, same 177
classes. --xml-pretty and --object-directory still work on 8.6 but are now
deprecated aliases for --cobertura-pretty and --gcov-object-directory, worth
knowing for whoever removes --object-directory next.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 16:35:16 -04:00
Frédéric Desbiens b6a00a2014 Added the Dependabot configuration the pinned actions need (#662)
The action references were pinned to commit SHAs in #660, and a SHA pin with
nothing moving it is worse than a floating tag -- it holds CI on whatever was
current the day it was written. That is exactly how actions/cache@v1 stayed in
ci_cortex_m.yml until GitHub began auto-failing every request that used it.
The drift measured before that catch-up: download-artifact four majors behind,
checkout and upload-artifact three each, cache and upload-pages-artifact two,
with nothing ever reporting it. This closes the loop, and the reference to
.github/dependabot.yml that #660 left in each workflow's pinning comment.

Weekly, github-actions only. Patch and minor are grouped into one pull request
because they are the routine traffic and a queue reviewed one item at a time is
a queue that gets ignored. Majors stay ungrouped, one each, because every
breaking change this repository has met in an action has been a major.

Two choices worth stating rather than leaving to be rediscovered.

target-branch is dev. Dependabot reads this file from the default branch, which
is master, but master is deliberately kept behind dev and pull requests belong
where the regression suites gate them. The consequence is that landing this on
dev arms it without firing it: nothing happens until a release merge carries
the file to master. Setting target-branch also opts out of Dependabot security
updates, which only run against the default branch -- a small cost for this
ecosystem, since an action advisory arrives as an ordinary bump on the weekly
run, but a real one.

The pull-request limit is raised from the default five to ten. Nine actions are
in use, and five would hold majors back with nothing saying that it had.

No other ecosystem is configured, deliberately: external dependencies are
forbidden, there are no submodules, and the one pinned tool -- gcovr in
scripts/install.sh -- lives in a shell script no ecosystem can parse, so that
pin keeps moving by hand.

No sibling eclipse-threadx repository has a Dependabot configuration, so this
sets the pattern rather than following one. The dependencies label it uses
already exists here.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 16:05:04 -04:00
Frédéric Desbiens 3d852eb451 Pinned every action to a commit SHA, and moved them off Node 20 (#660)
Node 20 is removed from the GitHub runners on 16 September 2026. Every run
in this repository currently emits the deprecation warning for it, naming
actions/checkout, actions/configure-pages, actions/upload-artifact,
LouisBrunner/checks-action and marocchino/sticky-pull-request-comment among
others. After that date those actions stop working rather than warning, so
this is a deadline and not housekeeping.

Every action is now referenced by a 40-character commit SHA with the version
in a trailing comment. A tag can be repointed at any commit; a SHA cannot, so
this is what makes "which code ran in our CI" answerable from the repository
rather than from whatever the tag meant at the time. The versions were behind
by as much as four majors -- download-artifact was on v4.3.0 against v8.0.1 --
because nothing in this repository has ever reported that an action moved.

Compatibility was checked against each new action.yml rather than assumed,
for every input this repository actually passes:

  checkout            submodules is unchanged
  cache               path and key are unchanged
  upload-artifact     name, path and retention-days are unchanged
  download-artifact   pattern, merge-multiple and path are unchanged
  configure-pages     takes no input here, and none became required
  deploy-pages        still exposes page_url, which the job reads
  upload-pages-art.   path is unchanged
  checks-action       token, name, conclusion, output and
                      output_text_description_file all survive v2 to v3
  sticky-comment      header and path survive v2 to v3, and the new
                      GITHUB_TOKEN input defaults to github.token, which is
                      what v2 used implicitly
  delete-artifact     name survives v5 to v6, and useGlob still defaults to
                      true, so the coverage_report-* glob from #655 still
                      matches
  CodeCoverageSummary already current at v1.3.0; pinned, not moved

The artifact pair moves together, as it must. The round trip was verified on
a runner before this commit: upload-artifact v7 to download-artifact v8,
through the pattern and merge-multiple selection #655 introduced, filtered 4
artifacts to 2 and produced exactly the tree the deploy expects.

Two behaviour changes worth knowing. download-artifact v8 adds a
digest-mismatch input defaulting to error, so a corrupted artifact now fails
the job instead of passing through -- the right default, but a change.
upload-artifact v6 and above require a runner of at least 2.327.1, which the
hosted runners satisfy and a self-hosted runner would need checking for.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 15:52:49 -04:00
Frédéric Desbiens 042049a7b2 Revived the Cortex-M build, which had compiled nothing since June (#653)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
Two defects, and the second hid the first.

This workflow triggered on master only, for both push and pull_request,
while dev is the integration branch. So it gated no pull request that
anybody opened. That is the same defect ports_arch_check.yml carries a
comment about, where it cost eight months of ports drifting from ports_arch
unnoticed, and regression_test.yml has it too.

And it had not compiled anything since at least 2026-06-08. Every run since
then failed in six to eight seconds at "Prepare all required actions",
before checkout, because GitHub automatically fails any request that uses
actions/cache@v1. The last run of any kind was 2026-06-30. A workflow that
fails in seven seconds is normally noticed within the day; this one was not,
because of the first defect. The two together meant the project's only job
that cross-compiles a port with GCC had been reporting nothing at all.

The toolchain now follows clang_check.yml rather than third party actions:
a pinned release fetched directly from Arm, verified against the published
sha256asc, and cached with actions/cache@v4. That also moves the compiler
off 9-2019-q4, a 2019 release, onto a version matching the GCC 14 default
this project states.

The ninja install is guarded on ninja being absent rather than run
unconditionally, because scripts/install.sh already carries a long comment
about apt-get update stalling for over two hours and taking a whole
regression run with it.

fail-fast is off so that one port failing still reports the other three.

Verified before committing, with the pinned 14.3.rel1 toolchain: the
download and sha256sum -c sequence in the install step was run as written,
the archive extracts to the directory the PATH step expects, and all four
ports configure and build clean with zero warnings.

This covers four ports of the forty under ports/ that have a gnu directory.
Widening it to every Arm gnu port is separate work.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 16:38:24 -04:00
Frédéric Desbiens 7959aef3bc Kept the coverage report from the runs that most need one (#659)
A failing test threw away coverage that had already been collected, and the
run whose behaviour changed is exactly the run whose coverage is worth
reading. Measured on the failing run of 2026-08-18: it uploaded test_reports
for all three suites and no coverage_report artifact at all.

Two causes, and the workflow one is the smaller of them.

cmake_bootstrap.sh runs under set -e, so a failing ctest aborted test()
before ./coverage.sh was reached. The gcda files exist by that point, so
nothing was missing except the step that reads them. ctest's status is now
captured and returned at the end, and the summary grep is allowed to fail
rather than being the thing that stops the coverage behind it.

The serial branch of the test dispatch collected no status either, so under
set -e the first failing configuration stopped the remaining four from being
tested at all -- and their coverage from being collected. That was cheap
while the suites ran in parallel, because the parallel branch already
collects exit codes from its background jobs. Moving to serial execution in
#643 quietly made one failure cost the other four configurations. The serial
branch now collects status the same way the parallel branch does.

With those fixed the report exists, so the workflow steps that publish it no
longer skip on failure. They are guarded with !cancelled() rather than
always(), so a cancelled run still stops promptly, which is the idiom
deploy_code_coverage already uses. The ${{ }} wrapping is required and not
decoration: a bare ! opens a YAML tag, and the file will not parse without
it.

Verified locally by replacing one test binary with a stub that exits 1:

  before  the failing run of 2026-08-18 produced no coverage_report artifact
  after   run.sh test default_build_coverage exits 8, and produces
          coverage_report/default_build_coverage.xml with 177 files and
          3804 of 3827 lines
  after   run.sh test all exits 8, and all five configurations run rather
          than stopping at the first

The failure still fails. Only the reporting around it changed.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 16:36:18 -04:00
Frédéric Desbiens adc6469b91 Published the coverage report instead of everything the run produced (#655)
The download step in deploy_code_coverage asked for the artifact named
${{ steps.artifact.outputs.coverage_report }}. That output is set by the
"Coverage Report name" step of run_tests, which is a different job, and the
steps context does not cross jobs. So the expression evaluated to the empty
string and the action took its documented path for an unspecified name:

    No input name, artifact-ids or pattern filtered specified,
    downloading all artifacts
    Total of 4 artifact(s) downloaded

The four are the two coverage reports and the two test_reports bundles of
JUnit XML, each extracted into a directory named after the artifact. The
next step uploads the lot to Pages, so the published site has carried the
test reports alongside the coverage, one directory deeper than intended,
under a path containing a run timestamp that changed on every publish. Any
link to a coverage report broke the next time one was published.

Selecting by pattern with merge-multiple fixes both halves: the pattern
excludes the test_reports bundles, and merging puts the contents of the two
coverage artifacts directly into coverage_report rather than under a
directory named for each. Each artifact holds one directory named for its
suite, renamed from default_build_coverage by "Prepare Coverage GitHub
Pages", so the result is the two suite directories the deploy expects and
the timestamped artifact name no longer appears in the published path.

Verified on a runner rather than reasoned about, with an isolated workflow
that uploads artifacts shaped like the real ones and downloads them both
ways:

    OLD  coverage_report/coverage_report-<epoch>-ThreadX/ThreadX/index.html
         coverage_report/coverage_report-<epoch>-ThreadX/default_build_coverage.xml
         coverage_report/coverage_report-<epoch>-SMP/SMP/index.html
         coverage_report/coverage_report-<epoch>-SMP/default_build_coverage.xml
         coverage_report/test_reports SMP/results.xml
         coverage_report/test_reports ThreadX/results.xml

    NEW  coverage_report/ThreadX/index.html
         coverage_report/SMP/index.html
         coverage_report/default_build_coverage.xml

Both artifacts carry a default_build_coverage.xml and the merge means one
overwrites the other, which the run above also shows. That file is consumed
by CodeCoverageSummary back in run_tests and is not read here, so it is
untidy rather than wrong, and it is called out in a comment.

The delete step is fixed in the same place and for a related reason. The
artifacts are named coverage_report-<epoch>, useGlob defaults to true in
this action, and as a glob "coverage_report" matches only the literal
string. It has been deleting nothing, without failing, and retention-days: 1
on the upload is what has actually been clearing these up.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 16:31:21 -04:00
Frédéric Desbiens eabdb86409 Matched gcov to the compiler that produced the coverage data (#658)
gcov reads a data format tied to the compiler that produced it. coverage.sh
took whatever gcov was first on PATH, which was fine while the compiler was
also whatever was first on PATH. #656 made cmake/linux.cmake honour CC and
#657 made a compiler switch actually reconfigure the build, so that
assumption no longer holds, and the first person to use the new capability
would have hit this.

Measured on dev with both of those merged:

    CC=gcc-14 ./run.sh build default_build_coverage    # succeeds
    ./coverage.sh default_build_coverage               # exit 64

gcov says why, if asked directly:

    tx_block_allocate.c.gcno:version 'B42*', prefer 'B33*'

gcovr turns that into "GCOV returncode was 3" and exits 64 through a Python
traceback, after the tests have already passed. It reads like a coverage bug
rather than a toolchain mismatch, which is the part that would have cost
someone an afternoon.

gcov is now derived from CC rather than found on PATH, so the caller sets
one variable instead of remembering two. GCOV still overrides, for a
toolchain that does not follow the gcc/gcov naming, and a derived gcov that
does not exist is reported as such instead of surfacing as a traceback.

Verified, tx and smp, before and after:

    CC=gcc-14    was exit 64, now exit 0, 177 files and 1527/3827 lines
    CC unset     exit 0, 177 files and 1527/3827 lines, unchanged
    CC=gcc-99    exit 1 naming gcov-99 and CC, rather than a traceback
    GCOV=gcov-14 with CC=gcc-99, exit 0, so the override still wins

A mismatched pairing still fails, deliberately: reading a gcc-14 tree with
the default gcc-13 gcov is exit 64 as before. Producing a number from
mismatched data would be worse than refusing.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 15:55:18 -04:00
Frédéric Desbiens a4397e133a Reconfigured the build when the requested compiler changes (#657)
CMake records the compiler it detected inside the build directory and keeps
using it on every later configure. Since cmake/linux.cmake began honouring CC,
that made a compiler switch silently ineffective: CC=gcc-14 ./run.sh build <cfg>
against an existing build directory printed "ninja: no work to do", exited 0,
and left the previous compiler in place. Anyone verifying a change against a
second compiler would have been reading stale results while being told the
build had succeeded.

generate() now compares the compiler recorded in the build directory with the
one currently requested, and reconfigures from scratch when they differ. The
comparison uses the path as CMake records it, which is the unresolved path as
given, so /usr/bin/gcc matches command -v gcc rather than the versioned target
its symlink points at. build_libs() gets the same treatment.

Only the C compiler is consulted: both trees that use this script declare
LANGUAGES C, so no CXX compiler is ever detected.

Nothing is reconfigured unless the compiler actually changed, so repeat builds
stay incremental and the default path is unchanged.

Assisted-by: Claude Code (Opus 5)
2026-08-24 14:47:26 -04:00
Frédéric Desbiens 5a68c9da4b Allowed the Linux toolchain file to accept a compiler override (#656)
cmake/linux.cmake set CMAKE_C_COMPILER and CMAKE_CXX_COMPILER unconditionally.
CMake reads a toolchain file before it consults CC and CXX, and a plain set()
in a toolchain file also takes precedence over -DCMAKE_C_COMPILER, so neither
of the two usual ways to pick a compiler had any effect: the tree could only
ever be built with whatever /usr/bin/gcc happened to point at.

That matters because AGENTS.md names GCC 14 as the project's default compiler
on Linux, while distributions still ship an older gcc as the default for some
time. Selecting GCC 14 previously meant either editing this file or changing
the machine's system-wide default.

Both variables now fall back to gcc and g++ only when nothing else has been
specified, so the default build is byte-for-byte what it was, while
-DCMAKE_C_COMPILER=gcc-14 or CC=gcc-14 now work as expected.

The binutils variables in this file are left alone: they are unused on the
Linux target, so guarding them would be unrelated churn.

Assisted-by: Claude Code (Opus 5)
2026-08-24 13:42:31 -04:00
Frédéric Desbiens 9218bad4bc Kept the coverage publish on master, where the environment allows it (#654)
Running the regression suites on dev (#652) was meant to test the branch the
pull requests target. It changed what gets published as well, which was not
intended and does not work: the first push to dev after that merge failed
with

    Branch "dev" is not allowed to deploy to github-pages due to
    environment protection rules.

All three suites passed in that run -- tx, smp and freertos. The only
failure was deploy / deploy_code_coverage, rejected before it ran, because
the github-pages environment restricts deployments to master.

The guard goes here rather than in the environment settings, because the
environment rule is doing its job. Which branch the published coverage
report describes is a deliberate decision, and moving it from master to dev
is a change worth making on purpose rather than as a side effect of a
trigger fix. Doing so needs the environment setting relaxed as well as this
line removed.

The per-suite deploy_code_coverage jobs need no guard: tx, smp and freertos
all pass skip_deploy: true, and regression_template.yml already restricts
that job to push and workflow_dispatch. Only the deploy job, which is the
one that publishes, was reaching the environment.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 13:07:38 -04:00
Frédéric Desbiens 977e14e776 Ran the regression suites on dev, where the pull requests actually are (#652)
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The ThreadX, SMP and FreeRTOS-compatibility suites triggered on master only,
for both push and pull_request. dev is the integration branch, so these
suites gated no pull request that anybody opened: the last dev run of any
kind was a manual workflow_dispatch on 2026-08-18.

This is the same defect ports_arch_check.yml already carries a comment
about, where it cost eight months of ports drifting from ports_arch
unnoticed. ci_cortex_m.yml has it too and is handled separately.

That 2026-08-18 run failed, which is the reason to check before switching
this on rather than after. Two tests failed: threadx_timer_simple_test in
the ThreadX suite, with ERROR #28, and threadx_thread_priority_change in
the SMP suite, with a timeout. Both were fixed two days later -- the first
by running the suites one test at a time (#643), which is what a timer test
failing only under parallel load wants, and the second by #647 by name.

Verified before this commit rather than assumed: both suites were re-run on
this tree, and all 1030 tests pass across all ten build configurations, the
ThreadX suite in 34 to 64 seconds per configuration and the SMP suite in 61
to 63. The 2026-08-18 run took 36m19s, of which a single test that has since
been given a budget accounted for 439 seconds.

No paths filter is added deliberately. The suites build the linux port, so a
filter would have to enumerate what cannot affect them, and the failure mode
of getting that list wrong is a regression that merges because the filter
excluded the file that caused it.

The deploy job needs no guard: regression_template.yml already restricts
deploy_code_coverage to push and workflow_dispatch, and restricts the
coverage PR comment to pull requests from the repository itself, so neither
fires for a pull request from a fork.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 11:30:19 -04:00