Commit Graph
100 Commits
Author SHA1 Message Date
Frédéric Desbiens 201cc04609 Applied R_ARM_RELATIVE relocations at GNU ThreadX module startup (#731)
Fixes #230

`gcc_setup.s` rebases the GOT and copies initialized data into the module data
area, but never applies the relocations the linker records for words inside that
data. A global initialized with a link-time constant -- `char *pBuffer = buffer;`,
a function pointer, a struct member holding a string literal's address -- kept
its link address and pointed outside the loaded module.

`gcc_setup.s` now walks `.rel.dyn` after the data copy and adjusts every
`R_ARM_RELATIVE` entry targeting the module data area, reusing the GOT loop's
base arithmetic and skip rules, and reloading the flash base first because
`crt0_memory_copy` clobbers `r3`.

The table is only emitted with `-pie`, so the example build scripts now pass
`-pie --no-dynamic-linker` and the example linker scripts discard `.dynamic`,
which `-pie` would otherwise place at address zero and inflate the image to
nearly 200 KB. Without `-pie` the table is empty and the loop is a no-op.

Covers Cortex-M3, M4, M7, M33 and M0+. Cortex-M23 ships no module linker or
build script and AArch64 uses a different relocation format; both left for a
follow-up.

Verified on `qemu-system-arm` with a module linked at one address and loaded at
another: char, function, string and struct-member pointers all resolve to the
loaded addresses, and the same module built without `-pie` still runs.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-15 16:24:54 -04:00
Frédéric Desbiens 3e5a16c95b Allowed tx_timer_change to be called from tx_application_define (#730)
Fixes #224

`_txe_timer_change` returned `TX_CALLER_ERROR` for any call at or above
`TX_INITIALIZE_IN_PROGRESS`, making `tx_timer_change` the only timer service
that could not be called from `tx_application_define` -- while still being
allowed from an ISR.

The check has no technical basis: `_tx_timer_change` only writes the expiration
fields of a timer that is not on an active list, with interrupts disabled.
Removed it from both the `common` and `common_smp` copies of
`txe_timer_change.c`, with the now-unused includes and the `TX_CALLER_ERROR`
line in the header comment. Relaxing an error check is backward compatible.

`testcontrol.c` in both suites now calls `tx_timer_change` at initialization, so
the timer simple test covers it: `ERROR #30` on `dev`, green with the fix.
116/116 SMP, 103/103 non-SMP.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-15 16:24:50 -04:00
Frédéric Desbiens 9b2979e6b0 Replaced the GNU-only dsb/isb 0xF operands with the UAL sy form (#729)
Fixes #551

`_tx_thread_system_return_inline()` in the Cortex-M `tx_port.h` headers spells
its barriers `dsb 0xF` and `isb 0xF`. A bare hexadecimal operand is a GNU
assembler extension, and IAR rejects it with `operand syntax error`, so the
header cannot be included at all. The block is guarded for GCC, armclang and IAR
together, so every IAR user of an affected port hits it -- four independent
reports on Cortex-M33 and M7 with EWARM 9.50 and 9.70.

Both operands become `sy`, the Arm UAL name for exactly what `0xF` encodes. The
generated instruction is unchanged. Applied to the two `ports_arch` masters and
all 32 copies under `ports`, covering M0, M23, M3, M33, M4, M52, M55, M7 and M85
across ac5, ac6, gnu, iar and keil, plus the `scripts/check_ports.sh` probes that
matched the old spelling.

`check_ports.sh` passes, the copy scripts still reproduce every generated port
byte for byte, and `arm-none-eabi-gcc -O2` compiles a caller for every patched
header, emitting `dsb sy` and `isb sy`. Three headers that need toolchain
intrinsics GCC does not ship fail identically on `dev`.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-15 16:24:45 -04:00
Frédéric Desbiens d15f28ab9f Masked the SMP remap core maps to silence a false -O2 array bounds error (#728)
Fixes #469

`_tx_thread_smp_remap_solution_find` indexes `_tx_thread_smp_schedule_list` with
the lowest set bit of the core maps it is given. Every caller already masks
those maps with `TX_THREAD_SMP_CORE_MASK`, but the compiler cannot see it, so
when the function is inlined into `_tx_thread_system_suspend` at `-O2` GCC
assumes a bit number as high as 31 and reports an out-of-bounds subscript.
`-Werror` turns that into a build failure.

Masked the three incoming maps at the top of the function, in both the inline
version in `common_smp/inc/tx_thread.h` and the twin in
`tx_thread_smp_utilities.c`. The masks are semantic no-ops, so scheduling is
unchanged, but the range is now visible to the optimizer. The first core queue
entry is also initialized, because the narrowed range lets GCC consider an empty
queue and warn about that instead.

Reproduced on the Cortex-A9 SMP port with arm-none-eabi-gcc 13.2.1, and clean
afterwards across -O2, -O3 and -Os and 2, 4 and 8 core configurations. SMP suite
116/116.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-15 16:24:40 -04:00
Frédéric Desbiens 1d4a4aeec8 Guarded _tx_thread_stack_analyze against inverted stack pointers (#727)
Fixes #460

`TX_ULONG_POINTER_DIF` casts the pointer difference to `ULONG`, so when
`tx_thread_stack_highest_ptr` sits below `tx_thread_stack_start` the midpoint
wraps to a huge value, the probe lands outside the stack, and the search never
converges: the caller hangs or faults.

`_tx_thread_stack_analyze` now requires the highest pointer to be strictly above
the start of the stack, and bounds the final scan by it. #464 covered the
`TX_THREAD_STACK_CHECK` path; this covers direct callers too, as
@billlamiework suggested on the issue.

Inconsistent pointers still mean a real overflow or a corrupted control block,
which remains the application's problem. What changes is that ThreadX reports it
through the stack error handler instead of hanging.

New cases for an inverted and an equal pointer pair crash the suite without the
fix and pass with it.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-15 16:24:35 -04:00
Frédéric Desbiens 5d235a534c Fixed the garbage _tx_initialize_unused_memory in the GNU Cortex-A ports (#726)
Fixes #435

`LDR x1, =__top_of_ram` already loads the top of RAM, so the `LDR x1, [x1]`
that followed read whatever sat at that address and left
`_tx_initialize_unused_memory` holding garbage.

Dropped that instruction from the 13 non-SMP GNU Cortex-A ports, the ARMv8-A
source they are generated from, and the two Cortex-A35 module examples. The SMP
GNU ports already had the correct form, and the Arm Compiler ports are
unaffected -- their symbol really does need the dereference.

`scripts/check_ports.sh` passes and the `gnu` CI job build-verifies the AArch64
ports. Not run on hardware.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-15 16:24:09 -04:00
Frédéric Desbiens cd6a2d9034 Put the module manager test's thread on the kernel's created list, so the stand-in kernel matches the one the manager reads (#739)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The host test built a TX_THREAD, gave it an ID and a module instance, and
left it off the kernel's created list. A created thread lives on that list,
so the thread the test created was not one the kernel would recognise.

The created list head and count are now defined for each object type, the
thread the test creates goes onto the thread list the way thread create puts
it there, and a successful delete takes it off again the way thread delete
does. Only the thread list is populated; the others stay empty because
nothing here creates an object of those types.

No expectation changed.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-14 17:09:00 -04:00
Frédéric Desbiens 83a3e62c28 Added a Module Manager host test harness and released a module thread's kernel stack only when it has one (#738)
* Released a module thread's kernel stack only when it has one, so deleting a thread without a manager-allocated stack no longer asks for that memory back

The thread delete dispatcher decided whether to release a kernel stack from
the calling module's property flags. A user-mode module can be given the
address of a thread that carries no kernel stack the manager allocated, and
deleting it passed that thread's null stack pointer to
_txm_module_manager_object_deallocate().

The pointer is now tested instead of the module's properties. Both thread
create paths clear the whole control block before filling it in, so a thread
with no manager-allocated kernel stack holds TX_NULL in the field, which
makes the test exact rather than an inference from how the module was
loaded.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Added a Module Manager host test harness and covered the user-mode thread kernel stack lifetime

The Module Manager is not in the ThreadX library the test tree links, and its
control blocks come from a module port rather than a base port, so nothing in
the suite could reach it. The thread_transition directory already describes
this technique as the one "the module manager tests in this tree use", but no
such tests existed.

This adds the directory that comment refers to: a port shim that replaces the
three interrupt primitives with host equivalents that count, and a CMake
target that compiles the manager sources under test directly against one
module port's headers.

The dispatch layer is one header of static functions and an unoptimised build
emits all of them, so a test that includes it to reach one dispatcher would
pull in references to every service the manager can dispatch. The header
guards each dispatcher with its own TXM_<SERVICE>_CALL_NOT_USED macro, so the
build reads that list out of the header and defines every guard except the
ones the test needs. Reading it rather than writing it down means a service
added later is excluded without anyone having to remember.

The first test asserts the invariant a user-mode module thread has to hold:
one logical thread costs the object pool two allocations, the control block
the module asks for and the kernel stack the manager takes on its behalf, and
deleting the thread must return the pool's available bytes and the module's
allocation-list count to exactly what they were. It asserts an equality
rather than a bound, because a bound would pass while one of the two was left
behind on every cycle, and then runs enough cycles that a leak of one stack
per cycle exhausts the pool several times over.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-14 16:08:48 -04:00
Frédéric Desbiens dde43b8ab2 Replaced the stale system stack switch pseudo-code in the ARMv7-A ports with comments that describe what the code actually does (#735)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The context save, vectored context save and system return routines in the
ARMv7-A ports carried pseudo-code comments claiming that they saved the
thread stack pointer and then switched to _tx_thread_system_stack_ptr.
Neither of those things happens, and none of these ports references that
variable outside an unused IMPORT in their example builds.

On ARMv7-A each processor mode has its own banked stack pointer. The IRQ
handler branches to _tx_thread_context_save while still in IRQ mode, so the
core's banked IRQ stack already serves as the system stack, and the thread
stack pointer is stored in the control block by _tx_thread_context_restore,
and only when the interrupt results in preemption. The scheduler runs on the
banked SVC mode stack that the startup code sets up. There is nothing for a
software stack switch to do.

The comments were therefore misleading rather than merely redundant, and had
led at least one user to try to restore the code they described. They are now
replaced by a description of the actual mechanism.

The AArch64 SMP ports keep their comments unchanged, because ARMv8-A does not
bank a stack pointer per processor mode and those ports do reload
_tx_thread_system_stack_ptr[core] explicitly.

This is a comment-only change. Every changed line is a comment, and all
twenty-five GNU variants still assemble cleanly for their target core.

The fourteen files under ports/cortex_a{5,7,8,9,12,15,17} were regenerated
from ports_arch/ARMv7-A/threadx/common/src/tx_thread_system_return.S with
ports_arch/ARMv7-A/update.sh. The ARMv7-A SMP ports have no generator, so
those files were edited directly.

Fixes #734

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-11 13:27:32 -04:00
Frédéric Desbiens c0aa4dbe29 Initialized the suspend status before the non-interruptable suspend in tx_thread_sleep, so a sleep no longer returns a stale error left over from an earlier timed-out suspension (#725)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
r52_fvp / r52 (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
When TX_NOT_INTERRUPTABLE is defined, _tx_thread_sleep did not set
tx_thread_suspend_status to TX_SUCCESS before calling
_tx_thread_system_ni_suspend. The interruptable path always did.

Nothing else writes that field on behalf of a sleeping thread:
_tx_thread_timeout resumes a TX_SLEEP thread directly and there is no
suspend cleanup routine for sleep. The value therefore survived from
whatever suspension the thread performed last, and tx_thread_sleep
returned it. A tx_semaphore_get that timed out with TX_NO_INSTANCE
would be followed by a tx_thread_sleep that also returned
TX_NO_INSTANCE after sleeping correctly.

The same change is applied to the SMP copy, which is identical.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-10 12:06:54 -04:00
Frédéric Desbiens 70741198e9 Refused thread delete and reset while an exit transition is in progress (#724)
_tx_thread_shell_entry and _tx_thread_terminate both publish a thread's
terminal state -- TX_COMPLETED or TX_TERMINATED -- and then call that
thread's exit notification callback, before the thread has been detached
from the ready list and before either service has finished with the pointer
it holds to the control block. That terminal state is exactly the state
_tx_thread_delete and _tx_thread_reset accept as authorization to
invalidate or rebuild the control block, and neither service tested whether
the transition producing it had finished.

A callback could therefore delete the thread it was called for -- and then
lawfully recreate it over the same memory, since delete exists to permit
that -- while the scheduler was still linked to the old incarnation.
tx_thread_create zeroes the whole control block and can auto-start the new
one, so the old priority list is left heading at a block whose own priority
field names a different list, with the old priority's map bit set behind
nothing. Alternatively a callback could reset a terminated thread, which
moves it out of the terminal state, and then resume it: the interrupted-
suspension logic in _tx_thread_system_resume refuses to void a suspension
only while the state is still terminal, so with the reset allowed first the
resume clears the suspending flag and restores TX_READY, the outer service
then finds the flag clear and skips the removal, and tx_thread_terminate
returns TX_SUCCESS for a thread that is runnable again.

Interrupt masking does not close the window, because the kernel restores
the prior posture before invoking the callback deliberately. On SMP
TX_RESTORE also releases the global protection, so the target can be
executing on another core while its callback runs -- and a reset there
memsets the stack a live core is running on. On the Linux, Win32 and Win64
host simulation ports the consequence is more immediate than corruption:
TX_THREAD_DELETE_PORT_COMPLETION cancels and joins the host thread backing
the deleted thread, so a callback-side delete of the completing thread
destroys the host thread the callback is running on.

The fix marks the transition and has the two services refuse a marked
target, returning the errors they already document, TX_DELETE_ERROR and
TX_NOT_DONE. The refusal is transient and the same call succeeds once the
transition has completed, so no documented lifecycle is lost; and it is in
the core services rather than the _txe_ wrappers, so disabling error
checking cannot disable it. Refusing the reset is also what closes the
resume path, without touching _tx_thread_system_resume: its existing
terminal-state test is sufficient once nothing can turn the terminal state
into TX_SUSPENDED from inside the window.

tx_thread_suspending is the marker, rather than a new control-block field.
It already means "a suspension is in progress" and is already true across
the callback in the two interruptable paths, so no field is added, the
public structure is unchanged, and sizeof(TX_THREAD) is unchanged --
which matters, because the Module Manager's object handling depends on the
sizes of the control blocks. Widening its lifetime was checked against
every reader rather than assumed. There are four: two in
_tx_thread_system_suspend and two in _tx_thread_system_resume. In every
window this change widens, the state is TX_COMPLETED or TX_TERMINATED, and
both resume readers already refuse to void a suspension for exactly those
two states, so their behaviour is unchanged; and no suspension routine is
called on the target in those windows, so the suspend readers never see
them. No suspension-initiating service can set the marker again inside a
window either: every one of them acts on a thread that is ready or
suspended.

Three sites needed changing beyond the two refusals, and the shape of each
was decided by where the marker can safely be cleared:

  - The non-ready branch of _tx_thread_terminate cleared the marker before
    the terminated extension and the callback, which is what left them free
    to act on a control block the service still had mutex-release
    processing to do against. The clear moves to the common tail, after the
    last dereference of the target, and becomes the single clear site for
    the whole service. In the interruptable ready branch the flag is
    already false there, because _tx_thread_system_suspend cleared it when
    it detached the thread, so the tail store is a second store of a value
    the flag already holds -- cheaper than testing for it, and it keeps one
    clear site.

  - Under TX_NOT_INTERRUPTABLE neither path set the marker at all, because
    that configuration does not use the interruptable suspension path that
    sets it. Both now set it before the callback. Interrupts being disabled
    there does not help: the callback is reached by a direct call.

  - In the TX_NOT_INTERRUPTABLE completion path the marker is cleared
    before _tx_thread_system_ni_suspend rather than after it. That call
    returns to the scheduler for a thread that is the current thread, which
    a completing thread is, and does not come back; clearing afterwards
    would leave a normally completed thread marked for ever and therefore
    permanently undeletable. Nothing is lost by clearing early there,
    because everything from that point to the detachment runs with
    interrupts disabled and calls no application code.

The change is the same change twice. All four files are byte-for-byte
identical between common and common_smp at this commit and stay so after
it, so common_smp was written by copying rather than by repeating the
edits. tx_thread_system_suspend.c and tx_thread_system_resume.c, which do
differ between the kernels, are deliberately untouched.

Tests. The in-tree regression test goes to both trees and is byte-for-byte
identical between them. It drives seven scenarios: terminating a ready
non-current target with two peers ready at the same priority, with the
callback attempting the delete and recreating the block if it succeeded;
the same with the callback attempting the reset and then the resume;
terminating a target suspended on a semaphore while owning a mutex, which
is the non-ready branch; natural completion alone at its priority,
including the safe post-completion reset, terminate, delete and recreate at
another priority; self termination; a benign callback, whose notification
count and ordering are unchanged; and the state and boundary cases, where
the new refusal must not fire.

It measures rather than describes. The callback records the published
state, the marker, and the status of every lifecycle service it can reach,
calling the core service as well as the wrapper wherever a refusal is
expected. A snapshot taken under interrupt lockout -- which is the global
SMP protection on an SMP port -- checks that every ready list agrees with
the control blocks it heads, that the priority map agrees with the lists,
and that each execute pointer is a member of the list its own priority
field names. Every walk is bounded, so a corrupted ring costs an assertion
and not a hang, and no test in the suite can hang. Expectations are counted
inside a scenario and gated between scenarios, so a failing kernel reports
how much it failed by without being driven further into its own
corruption.

The consequences are demonstrated from the terminator's context rather than
the completing thread's, which is what makes the pre-fix behaviour an
assertion instead of a wedged simulator. Compiled against the unfixed
sources the test fails 9 of the 24 expectations it reaches in the
uniprocessor tree and 10 of 24 in the SMP tree, and the failures are the
finding: the callback-side delete succeeds, the recreate succeeds, the
consistency snapshot disagrees, the target is neither terminal nor detached
when the service returns, and a peer has left the ready ring the recreated
block hijacked.

The SMP tree gets a second test for the case that needs concurrency. The
victim is excluded to core 1 and spins there without relinquishing while
the controller, excluded to core 0, terminates it, so the callback runs on
one core while the target executes on another. The callback-side reset and
delete must both be refused, and a sentinel written into the unused low end
of the victim's stack must survive -- a reset would have memset the whole
stack before rebuilding the frame. Every wait is bounded, and if the remote
precondition cannot be established the test says so and drops only the
assertions that depend on it rather than reporting a pass it did not earn;
measured over twenty consecutive runs it established the precondition every
time.

TX_NOT_INTERRUPTABLE and TX_DISABLE_ERROR_CHECKING are not among the five
build configurations either tree compiles, and each tree builds the whole
library once per configuration, so neither can be reached from inside the
suites. The lines this change adds under TX_NOT_INTERRUPTABLE are therefore
in no configuration the trees build, and they are where the permanent-
undeletability failure mode lives, so they get their own harness rather
than a compile check: the four sources plus the two error wrappers are
compiled directly into a test executable, once per combination, with
recorders standing behind the scheduler services they call. That is what
makes the marker's value at the moment of detachment directly observable.
It runs 66 expectations under TX_NOT_INTERRUPTABLE, 66 under that with
error checking disabled, 45 under that with notification disabled, and 63
under error checking disabled alone; against the unfixed sources those fail
25, 22, 6 and 20 respectively. The harness lives in the uniprocessor tree
only, because the four sources are identical between the kernels and the
shim replaces the very primitive the SMP port differs in, so a second copy
would compile the same text under the same macros. It is deliberately left
out of the coverage instrumentation, since the same source under different
feature macros has a different line set and merging those would confuse the
union rather than add to it.

Results. Both suites pass in all five configurations with GCC 14: 103 of
103 in the uniprocessor tree, up from 98, and 116 of 116 in the SMP tree,
up from 114. Merged line coverage is 100% in the uniprocessor tree and
5172 of 5183 in the SMP tree, whose eleven uncovered lines are the same
eleven that were uncovered before this change and are in tx_byte_pool_search
and tx_thread_smp_utilities; all four changed files are at 100% line
coverage in both trees, and SMP branch coverage rises from 2819 of 3548 to
2831 of 3556. Cross-compiled with arm-none-eabi-gcc at -Wall -Wextra for
Cortex-M4 against common and for Cortex-A7 SMP against common_smp, all four
files produce no diagnostics at all and an identical warning set to before
the change, under -std=gnu99 and -std=c99 alike -- unlike the module ports,
-std=c99 does not fail on these base ports, and even -Wconversion is clean.

No MISRA deviation is required: explicit comparisons to TX_TRUE, existing
ThreadX types, single-entry and single-exit control flow, no goto, and two
added constant-time tests that change no real-time complexity.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-10 12:04:40 -04:00
Frédéric Desbiens 9ee198a2d6 Made the stack analyze binary search require several consecutive fill words, so unwritten holes in a used stack no longer cause the highest stack pointer to be under-reported (#720)
The binary search in _tx_thread_stack_analyze() accepted a probe location as
unused as soon as a single word still held TX_STACK_FILL. A word inside an
otherwise used region that simply was never written - the padding of a
partially initialized local array, for example - therefore made the search
move away from the real boundary and report far less stack usage than the
thread had actually consumed, which in turn kept the stack guard from firing.

The probe now walks down from the candidate location and requires
TX_THREAD_STACK_ANALYZE_FILL_WORDS consecutive fill words before it treats
the location as unused, stopping early at the lowest location already known
to hold the fill pattern. The new macro defaults to eight words and can be
overridden in tx_port.h; setting it to one restores the previous behavior.
Holes shorter than the configured run no longer mislead the search, and the
result is unchanged for stacks that contain no such holes.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-10 11:02:42 -04:00
Frédéric Desbiens d56f96f37f Refreshed the reason gcc_check leaves RISC-V out (#722)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
r52_fvp / r52 (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The header explained the exclusion by saying RISC-V "is not regressing" and
that adding it would widen the toolchain download. The second half is still
true; the first read as though nothing in CI exercised the family at all,
which stopped being the case with #717.

RISC-V is now the best-covered of the four excluded families rather than the
least: regression_test.yml builds both ports and runs 955 tests on them under
QEMU -- 475 on RV32 and 480 on RV64, across five build configurations each --
which is more than a compile-and-link check could establish. That is a
stronger argument for leaving it out of this workflow than the original, so
the sentence now makes it.

The download figure is kept and quantified: the two bare-metal toolchains this
workflow would have to fetch are about 500 MB apiece.

Comment only; no behaviour change. scripts/check_gcc.sh names the same four
families but states the exclusion without giving a reason for it, so it needs
no matching edit.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-10 09:31:58 -04:00
Frédéric Desbiens 6c84e61d19 Gave install_riscv.sh the network hardening install.sh already had, and shared it between them (#721)
install.sh grew a retry loop, per-command timeouts and a deliberately
non-gating apt-get update after this runner pool cost several whole runs: a
mirror going silent for two hours, and a Hash Sum mismatch from a third-party
repository the project does not even use turning builds red. Those lessons were
local to that one file.

install_riscv.sh had none of them, and #717 puts it on every pull request's
critical path. Under set -e its bare apt-get update was a single point of
failure for the whole suite -- the precise case install.sh downgrades to a
warning on purpose -- and its two wget calls, each fetching about 500 MB, had
no retry and no timeout.

Rather than copy the helpers and let them drift again, they move to
tx_ci_common.sh and both scripts source it, following the arrangement
scripts/tx_windows_common.ps1 already uses on the Windows side. install.sh
keeps its behaviour exactly: same APT_OPTIONS, same 120-second TIMEOUT, same
three-attempt retry, and the comments explaining each of them travel with the
code they explain.

Two things are new:

  - TIMEOUT_LONG, 180 seconds, for a single large download. Sized against the
    39 seconds each tarball took on 10 Sep 2026 and deliberately not larger:
    the install step is capped at ten minutes, and a per-attempt timeout able
    to swallow that cap would leave the retry loop no turn to take, which is
    the failure mode the apt comment already records.

  - fetch(), which verifies a SHA-256 before anything is unpacked. Both digests
    were taken from the releases API and then checked against the bytes the CDN
    actually serves. This is not an independent trust root -- expected value and
    file come from the same host -- but it pins the bytes, so a deleted and
    re-pushed tag or a replaced asset stops the build instead of being picked up
    silently.

Also verifies qemu-system-riscv32 alongside riscv64. run.sh selects one per
architecture, so both are worth failing on here rather than at the first test.

Verified locally: retry returns 0 on success and 1 after three attempts;
fetch accepts a correct digest and, on a wrong one, fails and removes the
partial file; the source line resolves from the repository root, from an
absolute path and through a symlink; and both recorded digests match the
bytes served for the pinned tag.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-10 09:28:44 -04:00
Frédéric Desbiens 42f1d2bcf0 Pointed the SMP coverage comment at the pull request that made its measurements (#719)
The paragraph explaining why the SMP floor sits at 99 rather than 100 ended
with a bare reference that resolves to nothing a reader of this repository can
open. It named a source outside the tree in place of one inside it.

The real source is #677: it wrote the four tests that closed 53 of the 64
lines, took the measurements the paragraph quotes -- the 180,003 windows with
zero handovers, and the three remaining lines in tx_thread_smp_utilities.c --
and raised the floor from 98 to 99. Citing it gives the next reader somewhere
to go.

This is the same correction #718 made across five other files; this line was
missed because it sits in regression_test.yml rather than the template.

Comment only; no behaviour change.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-10 09:06:13 -04:00
Frédéric Desbiens 39277cf026 Stated the compiler and coverage requirements directly, and said who the pinned toolchain default serves (#718)
* Stated the compiler and coverage requirements instead of citing a file no contributor can open

Seven comments across five files cited a maintainer-local document as the
source for two project requirements: that GCC 14 on Linux is the default
compiler, and that the coverage target is 100%. That document is not part of
this repository and is not published anywhere, so the citation gave a reader
nothing to follow -- it named a source they cannot open, in place of simply
stating the requirement.

Both requirements are real and both stay. Only the pointer goes: each comment
now states the requirement on its own terms, which is what the surrounding
prose was already doing everywhere else.

  cmake/cortex_r52.cmake                     the pinned reference toolchain
  scripts/check_gcc.sh                       why the script exists
  .github/workflows/gcc_check.yml            why the workflow exists, and the
                                             GCC_VERSION pin
  .github/workflows/r52_fvp.yml              the GCC_VERSION pin
  .github/workflows/regression_template.yml  the coverage floor, twice

Comments only; no behaviour changes. Two paragraphs are rewrapped where the
shorter text left a ragged line. Verified that scripts/check_gcc.sh still
parses and prints its help from the header range it slices, and that
cmake/cortex_r52.cmake still configures the Cortex-R52 build.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Said who the pinned toolchain default in the CMake toolchain files serves

Both cmake/cortex_r52.cmake and cmake/cortex_m52.cmake default
ARM_TOOLCHAIN_PATH to a toolchains directory under the user's home, guarded by
an EXISTS check. Nothing said whether CI relies on that, and the natural
reading is that it does.

It does not. The three workflows that install a toolchain unpack it into the
workspace and cache it there, and r52_fvp.yml puts that directory on PATH
before configuring; scripts/check_gcc.sh passes -DARM_TOOLCHAIN_PATH at each of
its three CMake call sites. On a runner the guarded directory is absent, the
EXISTS check falls through, and the compiler comes from PATH. The default only
ever fires on a developer machine, where it is what makes a no-flag build work.

Both comments now say that, so the default is not mistaken for a CI dependency
and not removed as dead code. cortex_r52.cmake carries the explanation and
cortex_m52.cmake refers to it, matching the cross-reference already there.

The r52 comment also claimed absolute paths mean "the build does not depend on
PATH ordering", which is only true where the pinned directory exists -- in CI
the build depends on PATH and nothing else. Qualified accordingly.

Comments only; no behaviour changes. Verified that both toolchain files still
configure, and that the fall-through is real: with HOME pointed at a directory
holding no toolchains, cortex_r52.cmake configures against the arm-none-eabi-gcc
found on PATH.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-10 07:55:11 -04:00
Frédéric Desbiens 5e9d046b12 Compiled the module manager C sources with GCC, as clang already did (#716)
#689 added the module manager stage to check_clang.sh alone. The GCC half was
never written, so the module manager C stayed unbuilt by the project's declared
default compiler: 28 files of portable module manager under common_modules,
plus the three to nine per-port files under
ports_module/<core>/gnu/module_manager/src, across nine Arm module ports.

The stage is deliberately check_clang.sh's, port for port and header for
header, because a port covered by one check and not the other implies a parity
the checks list does not have. The same two details the ports dictate carry
over: an SMP port's control blocks come from common_smp rather than common, and
the TrustZone ports need -mcmse for their cmse_nonsecure_entry functions to be
honoured rather than ignored. The clang-only waiver does not: GCC implements
the optimize attribute that tx_thread_secure_stack.c carries, so nothing needs
suppressing for that file.

One divergence is forced by the toolchain. txm_module_manager_absolute_load.c
carries a #pragma message steering callers to the extended entry point, and the
C stages treat any compiler output as a failure. check_clang.sh silences it with
-Wno-#pragma-messages; GCC has no equivalent, and neither -Wno-pragmas nor any
other -W option suppresses the note -- verified with 14.3.rel1. The note is
therefore filtered out of the stage's output instead, together with the source
quote GCC prints beneath it. The filter stops at the next line that begins a
diagnostic of its own, so an error immediately following a waived note is still
reported; that case is what the injected-defect run below checks.

The workflow needed no trigger change: #689 added common_modules/** to both
path lists in advance, for the stage that had yet to arrive. Its header comment
is brought in line with what the workflow now runs.

Verified with the toolchain CI pins, arm-gnu-toolchain 14.3.rel1, both triples.
The full script passes and every port compiles every file, matching the counts
check_clang.sh reports for the same nine ports:

  cortex_a35 31/31   cortex_a35_smp 31/31   cortex_a7 34/34
  cortex_m0+ 33/33   cortex_m23 37/37       cortex_m3 33/33
  cortex_m33 37/37   cortex_m4 33/33        cortex_m7 33/33

Verified that the stage fails as intended by injecting defects into a throwaway
worktree: one in common_modules on the line straight after the waived pragma
note, reported under all nine ports, and one in a single port's own source,
reported only under that port.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-10 07:25:37 -04:00
Frédéric Desbiens 7f6e496233 Added the missing VFP enable field to the Cortex-R5/AC5 thread control block, which VFP builds were writing over the FileX pointer (#715)
The Cortex-R5/AC5 port defines TX_THREAD_EXTENSION_2 as empty while its
assembly reads and writes the per-thread VFP enable flag at [thread, #144].
With TX_ENABLE_VFP_SUPPORT that offset lands on tx_thread_filex_ptr, so
tx_thread_vfp_enable corrupted the FileX pointer and lazy save/restore
tested an unrelated value.

Defined TX_THREAD_EXTENSION_2 as ULONG tx_thread_vfp_enable, matching the
Armv7-A ports and the Cortex-R4/R5 GNU and AC6 ports fixed earlier, and
added tx_port_offset_check.c so a future layout change breaks the build
instead of silently retargeting the accesses. The offset was measured at
144 for this port and asserted.

Fixes #382

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-09 17:17:37 -04:00
Frédéric Desbiens 383dd311cd Removed the vector table offset register and system stack pointer setup from the Cortex-M low-level initialization (#714)
_tx_initialize_low_level wrote VTOR and derived _tx_thread_system_stack_ptr from
the reset vector on every Cortex-M port. Programming VTOR is the job of the
low-level startup code that runs before the kernel is entered, and doing it again
inside ThreadX silently overrode a vector table that the application, a
bootloader or a firmware update had already installed. It also forced every
application to export a vector table symbol that the kernel itself never needed.

_tx_thread_system_stack_ptr is never read by any Cortex-M port, so the value
copied out of the reset vector served no purpose.

Both blocks are removed from all Cortex-M ports, together with the now-unused
symbol declarations. The GreenHills files keep the NVIC base address load that
the removed block used to leave in r0 for the SysTick setup that follows.

Applications that relied on ThreadX programming VTOR must now set it in their
startup code. The bundled examples link their vector table at address zero and
run on the reset default.

Fixes #370

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-09 16:28:59 -04:00
Frédéric Desbiensandr 16297d7a2a Added a workflow that runs the Cortex-R52 images on the Armv8-R AEM FVP (#688)
* Added a workflow that runs the Cortex-R52 images on the Armv8-R AEM FVP

Nothing in CI executed a single instruction of any ThreadX port. gcc_check
compiles and links -- its own header says it "executes nothing" -- and its
CMake stage covers one R52 configuration, the base FVP example. A change that
assembles cleanly, links cleanly and then hangs on the first context switch
passed every required check.

This workflow builds both supported configurations and runs them on the
Armv8-R AEM FVP, asserting each image's self-reported result. The feature
configuration matters as much as the default one: VFP with the hard float ABI,
FIQ, IRQ nesting and FIQ nesting gate whole assembly blocks that the default
configuration never assembles, and it builds eight images where the default
builds five.

Arm distributes the model free of charge but behind a click-through licence
with no stable unauthenticated URL -- the developer.arm.com permalink forms
all 404 for it -- so the download location is a repository variable rather
than a literal. When it is unset the build lanes still run and still gate the
pull request; only execution is skipped, and it says so in the log and in the
job summary rather than passing quietly.

The image list is read from the generated ninja graph rather than kept in the
workflow. The images are EXCLUDE_FROM_ALL, so a bare build reports "no work to
do", and reading the graph means a target added to CMakeLists.txt cannot
escape the check by nobody remembering to list it here.

The module manager port and the S32Z280 targets are not on dev yet. Their
lanes belong in this workflow when they land, not in a second one.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Kept the FVP cache honest and stopped ctest passing on an empty suite

Two corrections to the new workflow.

The FVP cache key named the URL and hashed the workflow file. hashFiles()
reads files, so it cannot hash a repository variable, and pointing it at
this workflow keyed the cache on the file's own contents instead. That is
wrong in both directions: repointing FVP_AEMV8R_URL at a different model
kept serving the old one from cache, which is exactly what the comment
claimed the key prevented, and editing anything in this file forced a
needless refetch of a large archive. A small step hashes the URL and the
key uses that, so it now tracks what it names.

ctest reported success when it found nothing to run. Its default is to
pass an empty suite, and the example registers its tests inside both
if(FVP found) and if(Python3_FOUND). Either one going unmet on a runner
would have left this stage green having executed no image at all, which
is the hole the workflow exists to close. --no-tests=error makes an empty
suite a failure.

Verified with the pinned toolchain, Arm GNU 14.3.rel1: both configurations
configure, enumerate and build, and the image lists are the ones claimed.

  default: 5 images  boot_check demo_m2 demo_m3 demo_mpu demo_threadx
  feature: 8 images  the above plus demo_fiq demo_m5 demo_nesting

Also checked that the target enumeration survives the real ninja graph:
the generated graph offers twenty .elf paths, and the filter reduces them
to those eight, dropping the cmake_object_order_depends_target aliases and
the CMakeFiles copies.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

---------

Co-authored-by: r <r@r>
2026-09-09 14:39:54 -04:00
Frédéric Desbiensandr 8be4c45cce Compiled the module manager C sources, which no check had ever built (#689)
* Compiled the module manager C sources, which no check had ever built

#672 corrected this script's assembly glob and brought the module ports into
the count, but only their assembly. Their C stayed outside every check: 27
files of portable module manager under common_modules, plus the per-port code
under ports_module/<core>/gnu/module_manager/src. By this script's own standard
-- a port simply absent from the count reads as covered -- 293 files across
nine Arm module ports were compiled by nothing, with either compiler.

Each module port ships its own tx_port.h and txm_module_port.h carrying the
control-block extensions the dispatch code needs, so a port is compiled against
its own headers rather than the base port's. Two details the ports themselves
dictate:

  An SMP port's control blocks come from common_smp. Pairing cortex_a35_smp
  with the single-core headers hid _tx_thread_smp_protect and
  _tx_thread_smp_unprotect behind implicit declarations and lost
  tx_thread_smp_core_executing from TX_THREAD, so fourteen files reported
  errors for a port that builds correctly.

  The TrustZone ports carry cmse_nonsecure_entry, which needs -mcmse to be
  honoured rather than ignored. tx_thread_secure_stack.c also carries GCC's
  optimize attribute, which clang does not implement; that divergence is
  suppressed by name, for a file GCC builds cleanly.

All nine ports compile: 293 of 293. Verified that the stage fails as intended
by injecting a defect into a throwaway worktree -- a defect in common_modules
is reported under every port, one in a port's own source only under that port.

Assisted-by: Claude Code (Opus 5)

* Ran the module manager stage against the sources that reach it

Three corrections to the new stage, all found by running it against dev
rather than against the tree it was written on.

The workflows did not trigger on common_modules. Both check_clang.yml and
check_gcc.yml list ports_module but not common_modules, and the portable
module manager under it is the larger half of what the stage compiles --
28 of the 31 to 37 files each port builds, plus two include directories.
A change there left the stage unrun, which is the same "absent from the
count reads as covered" that the stage exists to close. Both lists gain
common_modules, and they stay identical to each other as the comment in
each asks.

A deliberate deprecation notice read as a build failure. Since this PR
was opened, txm_module_manager_absolute_load.c gained a #pragma message
steering callers to the extended entry point. The stage treats any
compiler output as a failure, so that one notice failed every port: nine
failures on a tree where nothing is wrong. Pragma messages are now
waived for the stage, because a notice to callers is not a defect in the
file that carries it.

That failure also printed nothing. Both C stages report by grepping the
output for "error:", so a diagnostic that is not an error produced a bare
FAIL line with no reason under it, and the only way to learn the reason
was to reproduce the compile by hand. Both stages now fall back to
showing what the compiler actually said.

The TrustZone attribute waiver is narrowed to the one file that needs it.
tx_thread_secure_stack.c carries GCC's optimize attribute, which clang
does not implement; it is the only file among the 300-odd this stage
compiles that does. Waiving the warning for the whole port would have
swallowed a stray unknown attribute anywhere else in it.

Verified with the same toolchain CI uses, ATfE 22.1.0: the full script
passes, and every port compiles every file.

  cortex_a35 31/31   cortex_a35_smp 31/31   cortex_a7 34/34
  cortex_m0+ 33/33   cortex_m23 37/37       cortex_m3 33/33
  cortex_m33 37/37   cortex_m4 33/33        cortex_m7 33/33

The counts are each one higher than this PR first reported, because
txm_module_manager_absolute_load_extended.c has landed since. The stage
picked it up with no edit, which is what globbing the directories was for.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

---------

Co-authored-by: r <r@r>
2026-09-09 14:39:30 -04:00
Frédéric Desbiens 6ce8d5cc76 Marked every published ThreadX include directory as SYSTEM so applications no longer get warnings from ThreadX headers (#713)
* Marked every published ThreadX include directory as SYSTEM so applications no longer get warnings from ThreadX headers

Commit 8c3c08f added the SYSTEM keyword to the target_include_directories
call in common/CMakeLists.txt, but the same call in every port, in the
SMP common directory's consumers, in the POSIX and FreeRTOS compatibility
layers, and for the generated tx_user.h directory was left unchanged.
A consumer building with a strict warning set therefore still saw
diagnostics coming from tx_port.h and from the compatibility layer
headers, which is exactly what the original change set out to avoid.

Added the SYSTEM keyword to all of those calls so CMake emits -isystem
rather than -I for every directory holding a ThreadX public header. The
ARMv7-M, ARMv8-M, ARMv7-A and ARMv8-A architecture sources under
ports_arch were updated alongside the ports they generate, keeping the
two in step.

Directories that are PRIVATE to an example or test build were left as
they are, since nothing is published from them.

Fixes #290

Assisted-by: Copilot (Opus 5) <noreply@github.com>

* Stopped apt-get update being a gate it was never meant to be

apt-get update fails if any configured repository serves a bad index,
including ones this project never reads. The GitHub runner image carries
Google's and Microsoft's apt repositories, and a Hash Sum mismatch from
Google's, their CDN caught mid-publish with the index and the Release
file eight hours apart, failed all three attempts and turned a run red
over a browser nobody was installing.

Made a failed update warn and carry on, leaving apt-get install as the
gate. Nothing is weakened by that: the install still exits on a package
it cannot find, so an unreachable archive still stops the script, one
step later and naming the package it could not get, which is a better
diagnostic than a hash mismatch in a repository nobody asked for.

Disabling third-party sources before updating would keep the update
strict, but this script also runs on a contributor's own machine, and
rewriting someone's apt configuration to suit CI would be worse than
tolerating a stale index for an archive we do not read.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-09 14:27:02 -04:00
Frédéric Desbiens 3fc28d8979 Allowed a BASEPRI-masked interrupt to wake the Cortex-M idle loop from WFI (#711)
The idle loop in _tx_thread_schedule masks interrupts before re-reading
_tx_thread_execute_ptr, so that an interrupt cannot make a thread ready
between that read and the WFI and then be lost. With PRIMASK that is safe:
the architecture excludes PRIMASK from the WFI wake-up condition, so the
interrupt still wakes the processor and is taken as soon as the loop
re-enables interrupts.

BASEPRI is not excluded. When TX_PORT_USE_BASEPRI is defined, the mask
written at the top of the loop is still in place at the WFI, so any
interrupt at or below TX_PORT_BASEPRI fails to wake the processor at all.
On a system where every interrupt is managed by ThreadX, that is every
interrupt, and the processor stays in WFI or in the low power mode entered
by tx_low_power_enter until something outside the mask happens.

Set PRIMASK and clear BASEPRI for the duration of the WFI, then restore
BASEPRI and clear PRIMASK. Interrupts remain masked across the whole
window, so the original race is still closed, but the wake-up condition is
now evaluated with BASEPRI clear and any enabled interrupt can end the wait.
A pending interrupt is taken at the CPSIE i, or at the existing unmask
further down the loop if it falls below TX_PORT_BASEPRI.

The change is made in the ARMv7-M and ARMv8-M sources under ports_arch,
including the module manager, and propagated to the generated ports with
the copy scripts. The ghs and keil ports do not implement
TX_PORT_USE_BASEPRI, and Cortex-M0/M0+/M23 have no BASEPRI register, so
those ports are unaffected.

Verified by assembling every patched GNU port for Cortex-M3, M4, M7, M33,
M55 and M85, with and without TX_PORT_USE_BASEPRI, TX_LOW_POWER and
TX_ENABLE_EXECUTION_CHANGE_NOTIFY, and by checking the disassembly of the
idle loop.

Fixes #279

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-09 11:29:14 -04:00
Frédéric Desbiens 146d57b235 Fixed the RISC-V64 trap frame size mismatch in the regression test BSP (#708)
The RISC-V64 port moved its interrupt frame to 528 bytes (65 slots plus 8
bytes of padding, so sp stays 16-byte aligned at a call) and published the
size as TX_RISCV_TRAP_FRAME_SIZE. The port sources were converted to use
it, but the shared regression test BSP was not: its trap_entry still
allocated a hardcoded 65 * REGBYTES, or 520 bytes.

Every interrupt therefore unwound 8 bytes more than it allocated:

    trap_entry:                   addi  sp,sp,-520
    _tx_thread_context_restore:   addi  sp,sp,528

On the RISC-V64 regression suite that left 24 of 95 tests failing in the
default configuration, typically as an illegal instruction once execution
reached a corrupted frame. The example BSP under the port directory was
converted with the port and was unaffected, which is why the functional
QEMU test kept passing.

The test BSP now takes both frame sizes from the port it is linked
against, so the two cannot drift apart again. A port that publishes no
contract keeps the historical layout, so the RISC-V32 side is unchanged
until its own port publishes one.

TX_RISCV_TRAP_CALL_FRAME_SIZE is restored to the RISC-V64 tx_port.h. It
was removed as unused when the frame sizes were introduced, but it is
part of the same contract: it is the space a trap entry reserves around a
call into C, and the psABI requires 16 bytes there rather than one
register slot.

Verified on QEMU with every linkable test built, comparing against the
commit before the port change:

    before the port change   2 failures out of 95 (both unlinkable)
    current dev              24 failures out of 95
    with this change          2 failures out of 95 (both unlinkable)

The two remaining failures predate all of this: newlib pulls _impure_ptr
out of R_RISCV_HI20 range for time(), so those two binaries do not link.
RISC-V32 is unchanged at 2 failures across all five configurations.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-09 10:56:57 -04:00
Frédéric Desbiens 9a03838381 Stopped the test runners from reporting an incomplete build as test failures (#709)
The cmake test runners drove Ninja with its default keep-going of 1, so the
first failing target ended the build. Every target scheduled after it was
simply absent, and ctest reports a missing binary as a failing test. A
single link error therefore produced a failure count that moved with build
scheduling order rather than with the code.

Measured on the RISC-V64 regression suite, where two targets genuinely
cannot link:

    before   79 of 95 test binaries built
    after    93 of 95 test binaries built

Fourteen perfectly good binaries were being skipped and counted as
failures. Passing -k 0 lets Ninja finish everything it can; the build still
exits non-zero when a target fails.

Three related problems in the same paths are fixed with it.

A failing configuration used to abort the loop over configurations, so
under set -e the ones after it went unbuilt or untested. The build loops
and the serial test loops now accumulate status and return it at the end,
which is what the parallel test branch already did with wait, and what
cmake_bootstrap.sh already documented for ctest.

Capturing that status removes the set -e protection inside the functions,
so two latent faults become reachable and are closed here. A failed pushd
would have let ctest run in the source tree, where it finds no tests and
reports success; the pushd is now guarded. And ctest's status was
discarded by the popd that follows it, so a configuration with failing
tests returned 0 and was reported as a pass; the status is now carried
past the popd and the summary steps.

Verified on the RISC-V32 suite, which has two genuine failures in each of
its five configurations. Both the serial and the parallel branch now test
all five and exit 8, where the serial branch previously stopped after the
first configuration.

The tx and smp runners are symlinks to scripts/cmake_bootstrap.sh, so they
are covered by the one change there.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-09 10:56:28 -04:00
Frédéric Desbiens 37f48fa83c Removed FIFO queueing from the ARMv7-A SMP ports so an ISR can no longer deadlock waiting for protection (#707)
The Cortex-A5, A7 and A9 SMP ports guarded the inter-core protection with
a FIFO wait list: a core that could not take the lock added itself to a
queue, incremented _tx_thread_smp_protect_wait_counts[core], and only the
core at the head of the queue was allowed to acquire.

Two parts of that scheme require the waiting core to be interruptible.
_tx_thread_smp_unprotect() refuses to release the protection while the
releasing core's own wait count is non-zero, and a queued core is taken
back out of the list by _tx_thread_context_restore() when it is preempted.
Neither can happen on a core that is spinning inside an ISR with
interrupts already masked, because the spin loop restores the caller's
interrupt posture rather than enabling interrupts. The core stays in the
list forever and the system deadlocks with cores stuck in the wait loop.

This is the deadlock Microsoft removed from the ARMv8-A SMP ports in
6.1.11, by dropping the wait list and using a plain LDAXR/STXR spinlock.
The same removal was announced for the ARMv7-A ports at the time but was
never made, so those three ports have carried the deadlock ever since.

The FIFO queueing is now removed from the ARMv7-A SMP ports as well.
_tx_thread_smp_protect() becomes an LDREX/STREX spinlock that releases
interrupts between attempts, matching the sequence already used by the
Cortex-R8 SMP port; _tx_thread_smp_unprotect() no longer consults the
wait counts; and _tx_thread_context_restore() no longer has to unqueue a
preempted core. The now unreferenced wait list macro headers are deleted.

Verified by assembling all three GNU ports with arm-none-eabi-gcc for
cortex-a5, cortex-a7 and cortex-a9, with and without
TX_ENABLE_FIQ_SUPPORT, TX_ENABLE_WFE and TX_MPCORE_DEBUG_ENABLE, and by
reading back the disassembly of the new protect sequence. The port
consistency checks pass.

Fixes #219

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-09 10:29:51 -04:00
Frédéric Desbiens d3fb72b6dd Aligned the simulator ports' fake stack pointer so ThreadX no longer performs misaligned ULONG accesses (#705)
The Linux, Win32 and Win64 simulation ports build a fake initial stack
pointer by subtracting a fixed 8 bytes from tx_thread_stack_end. That
field addresses the last byte of the thread's stack area, so it is one
less than an aligned address and the resulting pointer is misaligned by
construction, no matter how well aligned the stack the application
supplied was.

Two ULONG accesses then use that pointer. _tx_thread_stack_build() itself
clears the word below it, and _tx_thread_create() copies it into
tx_thread_stack_highest_ptr, which TX_THREAD_STACK_CHECK dereferences on
every suspend and resume when TX_ENABLE_STACK_CHECKING is defined. Both
are undefined behaviour. They happen to work on x86 but are reported by
GCC's undefined behaviour sanitizer, and would fault on a host that
requires natural alignment.

The fake stack pointer is now rounded down to a ULONG boundary, which
leaves it inside the stack area and makes both accesses aligned.

Verified by building the Linux port with -fsanitize=undefined and
running the demo: the two reported diagnostics are produced before the
change and neither appears after it. The tx and smp regression suites
pass, 98 and 114 tests respectively.

Fixes #218

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-09 09:34:33 -04:00
Frédéric Desbiens 2e48ce0d1b Removed the stray build artifacts committed with the RISC-V64 port fix (#706)
A local CMake build tree was staged by mistake alongside the RISC-V64
spec compliance work in #698, putting nine generated files on dev,
including a compiled kernel.elf, build.ninja, the Ninja dependency logs
and a QEMU run log. None of it belongs in the repository.

The root .gitignore listed build directories by name rather than by
pattern, so build/, build_qemu/, build_m7/ and the build_r52 variants
were covered but a differently named tree was not. Those entries are
replaced with a single build*/ pattern, which covers every existing name
and any future one. No tracked file matches the new pattern.

The artifacts remain reachable in history; only the working tree is
corrected, since rewriting a shared branch is the greater harm.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-09 09:31:00 -04:00
Frédéric Desbiens 515ab8aba1 Added the missing memory barrier so ARMv7-A SMP schedulers no longer miss preemptions (#704)
_tx_thread_schedule stores the newly selected thread into
_tx_thread_current_ptr[core] and then reloads _tx_thread_execute_ptr[core]
to detect a concurrent scheduling decision made by another core. On the
other side, _tx_thread_smp_core_interrupt stores the execute pointer and
then reads the current pointer to decide whether an inter-core interrupt is
required. This is a store-buffer pattern: without a barrier both cores can
observe stale values, the interrupt is skipped, and a ready thread with the
highest priority is never scheduled.

The barrier was added to the ARMv8-A SMP scheduler in 6.2.1, but the
ARMv7-A SMP ports were left untouched even though they implement the same
protocol. This adds the corresponding DMB to the Cortex-A5, Cortex-A7 and
Cortex-A9 SMP schedulers for both the AC5 and GNU toolchains.

The GNU variants were verified by assembling them with arm-none-eabi-gcc
13.2.1 for their respective cores.

Refs #209

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-09 09:10:36 -04:00
Frédéric Desbiens e99f0d4207 Corrected the module kernel stack size so it no longer overstates the usable stack (#701)
The module manager recorded tx_thread_module_kernel_stack_size as the raw
TXM_MODULE_KERNEL_STACK_SIZE constant, but the end of the kernel stack is aligned
downwards to an eight-byte boundary while _txm_module_manager_object_allocate only
guarantees ULONG alignment. The recorded size could therefore overstate the usable
stack by up to seven bytes. The scheduler copies this value into tx_thread_stack_size
whenever a user mode module thread enters the kernel, so the overstated value is
visible to RTOS-aware debuggers and to anything built on it.

The size is now derived from the aligned end minus the start. Also documented that
TX_ENABLE_STACK_CHECKING is not supported for module threads.

Refs #181

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-08 16:44:48 -04:00
Frédéric Desbiens 945f5f5caa Moved the ARC ISR enter callout onto the system stack (#700)
_tx_thread_context_save() calls _tx_execution_isr_enter when
TX_ENABLE_EXECUTION_CHANGE_NOTIFY is defined. On the path where an interrupt
preempted a running thread, that call was made before the switch to the system
stack, so the 32-byte call frame and the whole stack footprint of the callout
were taken from the interrupted thread's stack, on top of the 160-byte interrupt
frame the port had already allocated there.

The callout is supplied by the application, so its stack usage is not bounded by
ThreadX, and it is charged to every thread that happens to be running when an
interrupt arrives.

The two other callout sites in the same routine, the nested save and the idle
system save, already run on the system stack, as does the _tx_execution_isr_exit
call in _tx_thread_context_restore. _tx_thread_schedule was reordered in 6.1.9
so that _tx_execution_thread_enter runs on the system stack rather than the
thread stack; the same reorder was never applied to the context save.

The switch to the system stack now happens before the callout in the ARCv2_EM,
ARC_HS and SMP ARC_HS ports. _tx_thread_context_fast_save is unchanged because
the fast interrupt path never switches stacks by design.

Fixes #149

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-08 15:42:39 -04:00
Frédéric Desbiens b0ec8bfbb9 Fixed the invalid module data pointers in the absolute module load (#699)
An absolutely located module has its code and its data placed at two
independent fixed addresses by the module's linker script. The module
preamble carries the code and data sizes but not the data address, so
_txm_module_manager_absolute_load() could not determine where the
module's data area was. It computed txm_module_instance_data_start
from the code size and the preamble size, which yields a size rather
than an address, and it set txm_module_instance_module_data_base_address
one past the end of the byte pool allocation.

Added _txm_module_manager_absolute_load_extended(), which accepts the
module's data area address from the caller. Deprecated
_txm_module_manager_absolute_load(), which now forwards to the extended
service with an unknown data area location and rejects modules that
request memory protection, since the memory protection hardware cannot
be programmed to cover an unknown data area.

Fixes #450

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-08 10:03:45 -04:00
Frédéric Desbiens d7789f0b12 Fixed the clobbered return address in the RISC-V context save (#696)
_tx_thread_context_save() returns to its caller with ret, which uses the
return address held in ra. When TX_ENABLE_EXECUTION_CHANGE_NOTIFY was
defined, the call to _tx_execution_isr_enter overwrote ra with the address
of the instruction following the call, so the subsequent ret returned into
_tx_thread_context_save itself instead of the interrupt service routine.

The return address is now saved on the stack around the call and recovered
afterwards, which is the same idiom already used by the Arm ports. The fix
covers all three affected paths (nested save, thread save and idle system
save) in the risc-v32 GNU, risc-v32 IAR and risc-v64 GNU ports.

Fixes #348

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-03 14:17:49 -04:00
Frédéric Desbiens 508af549da Fixed the incorrect loop bound constant in the IAR file lock support (#695)
The IAR multithreaded library support code allocates its file lock
mutexes from an array of _MAX_FLOCK entries, but the wrap-around check
and the exhaustion check in __iar_file_Mtxinit() both compared against
_MAX_LOCK, the bound of the unrelated system lock mutex array.

When _MAX_FLOCK is greater than _MAX_LOCK, the free mutex index wrapped
early and the exhaustion check reported failure while free entries
remained, so *m was set to TX_NULL and the application faulted the first
time a file lock was taken.  When _MAX_FLOCK is smaller than _MAX_LOCK,
the free mutex index was allowed to run past the end of
__tx_iar_file_lock_mutexes and the exhaustion check could never fire.

Corrected all four comparisons in each of the 27 copies of tx_iar.c.

Fixes #444

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-03 13:50:31 -04:00
Frédéric Desbiens 17ff21a07f Fixed the garbled comment in the POSIX condition variable sources (#694)
The comment above the internal semaphore lookup in the POSIX condition
variable implementation was garbled: it contained a capitalisation typo
("COndition") and an incomplete sentence ("into a semaphore  a cast").
The four affected sources now share a single, correct wording.

This is a comment-only change; there is no functional impact.

Fixes #419

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-03 12:58:26 -04:00
Frédéric Desbiens c469a7756b Fixed the missing immediate prefix on MOV in Cortex-M schedulers (#693)
The BASEPRI-masking path in tx_thread_schedule wrote "MOV r0, 0" rather
than "MOV r0, #0". UAL requires the "#" prefix on an immediate operand.
GNU as and the LLVM-based assemblers accept the unprefixed form and emit
the intended encoding, but stricter assemblers reject it outright, so the
affected ports could not be built with those toolchains.

The GNU and AC6 sources had already been corrected; this brings the IAR
and AC5 sources into line. Verified that both spellings assemble to the
same Thumb-2 encoding (f04f 0000), so this is a source-correctness fix
with no change in generated code or runtime behaviour.

Covers 31 occurrences across the Cortex-M3, M4, M7, M33, M52, M55 and M85
ports, their module manager counterparts, and the shared ARMv7-M and
ARMv8-M architecture sources.

Fixes #461

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-03 11:44:27 -04:00
Frédéric Desbiens 13c8c768c7 Added the boot-at-EL1 option to the S32Z280 entry path (#690)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The Armv8-R AEM FVP entry path has carried TX_R52_BOOT_AT_EL1 since it
was written, for the case its own comment describes: "an earlier boot
stage or a vendor EL2 monitor has already dropped privilege to EL1".
This board's entry path did not, so a kernel could not be built as a
guest on it at all -- and it is the board where that matters most,
because it is the one with silicon behind it.

The bracket is the whole change.  Everything from the Thumb reset
trampoline to the ERET goes inside #ifndef TX_R52_BOOT_AT_EL1, and the
#else supplies a one-instruction A32 _start that branches to el1_entry.

A32 AND NOT T32, which is the one real difference from the standalone
entry.  The core resets in Thumb state here because the RTU boot
instruction NXP plants is a T32 branch, but a guest is not reached by
reset: it is reached by the monitor's ERET, and the monitor chooses the
state through SPSR.T.  Get the two out of agreement and the guest dies
on its first instruction with an undefined-instruction exception, which
looks exactly like a bad entry address and sends the reader to the
loader instead of to the ERET.

WHAT THE MONITOR INHERITS is enumerated at the #ifndef, next to the code
it replaces rather than in a document, because that is where somebody
adding a third board will be looking.  This board's EL2 block is
considerably larger than the model's, and each item on the list is
something a guest at EL1 provably cannot do rather than something it
merely does not: CNTFRQ is writable only at the highest implemented
exception level and reads zero out of reset; HCPTR.TCP10/TCP11 reset
set, trapping every EL1 floating-point access; HSCTLR.TE is an EL2
register (SCTLR.TE is EL1's, and el1_entry still clears it below);
ICC_HSRE.SRE makes every other ICC_* and ICH_* register exist at all;
the low-latency peripheral port enables reset to zero and an EL1 write
to that register traps to EL2; and the TCM enables are per-core with
ENABLEEL2 SILENTLY IGNORED from EL1 -- measured on both BTCM and CTCM,
the base took and bit 0 took while bit 1 stayed clear.

CNTHCTL.PL1PCTEN and PL1PCEN are the deliberate omission from that list,
and the note says why.  This path opens both, because a standalone
kernel owns the physical timer.  A monitor that TIME-partitions its
guests must not: a partition's physical time keeps running while it is
descheduled, so a guest reading it can observe that it was not running.
That is the monitor's decision rather than this file's, which is why the
list says what a guest cannot do rather than what a monitor should.

Verified both ways.  The three standalone images build and link
unchanged, and a kernel built with the option boots at EL1 on a
S32Z280-594EVB under an EL2 monitor, runs two threads through a queue
and a semaphore, and reports back -- with no other change to the kernel
or to its port.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-02 19:08:24 -04:00
Frédéric Desbiens 8c681c188e Asserted that the Cortex-R52 port refuses the options it documents as refused (#687)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
ports/cortex_r52/gnu/CMakeLists.txt rejects two option combinations at
configure time -- TX_R52_ENABLE_VFP without TX_R52_FLOAT_ABI=hard, and
TX_R52_ENABLE_FIQ_NESTING without TX_R52_ENABLE_FIQ.  Nothing exercised
either.  A guard that has stopped firing looks exactly like a guard nobody
has tripped, so both could have been silently disabled by a typo in a
variable name at any point and no check would have noticed.

check_gcc.sh gains a sixth stage that configures each rejected combination
and asserts the configure fails.  It needs no new harness: the script already
runs cmake as a subprocess for the CMake example builds, and the port's
CMakeLists is included by the toolchain file alone, so no example flags are
needed and the two negative configures stop almost immediately.

Two things the stage does that a thinner version would not.

It asserts the message text, not just the exit status.  A configure that
fails for an unrelated reason would otherwise be recorded as a guard doing
its job.

It also configures the SUPPORTED combination and requires that to succeed.
Two negative assertions on their own are satisfied by a guard that rejects
everything -- the port would be unbuildable and the check would still pass.
The positive case is what separates a guard that works from one that is
merely always on.

Verified negatively, four deliberate breaks, each caught:

  VFP guard condition forced false      -> "configure succeeded, but this
                                           combination cannot build"
  FIQ nesting guard forced false        -> the same, on that case
  guard fires, message text changed     -> "configure failed, but not on the
                                           expected guard"
  VFP guard condition forced true       -> "the supported combination was
                                           refused"

The fourth is the one the positive case exists for and the only one a
refusals-only stage would have missed.  ports/cortex_r52/gnu/CMakeLists.txt
was confirmed byte-identical to dev afterwards.

Full run passes: exit 0, all six stages, and --asm-only correctly skips the
new one.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-01 09:42:38 -04:00
Frédéric Desbiens 57390fc0fe Refused the Cortex-R52 VFP option without a hard float ABI (#686)
TX_R52_ENABLE_VFP with the default soft float ABI cannot build.  The option
defines TX_ENABLE_VFP_SUPPORT, which enables the VMRS, VSTMDB and VLDMIA
blocks in the port assembly, and -mfloat-abi=soft leaves the assembler with
no FPU to accept them.  The configuration fails with eight errors of the form

  tx_thread_system_return.S:118: Error: selected processor does not support
  `vmrs r4,FPSCR' in ARM mode

none of which mentions the float ABI, so a user has to reason from VMRS back
to the option that enabled it.

The guard that exists said otherwise.  It warned that "the compiler will not
emit floating-point instructions, so the VFP context path will never be
exercised", which describes a build that succeeds and is merely pointless --
and then let configure finish, so the warning scrolled past well before the
assembler errors appeared.

It is now a FATAL_ERROR that names the fix, which is what the same file
already does six lines above for TX_R52_ENABLE_FIQ_NESTING without
TX_R52_ENABLE_FIQ.  That combination is rejected for being "meaningless",
while this one, which cannot assemble at all, was only warned about.  The
severities were the wrong way round.  The option's definition also moves
below the check, so the block reads like the FIQ nesting one.

The ABI is not promoted to hard automatically.  TX_R52_FLOAT_ABI is a cache
variable the user may have set deliberately, and silently overriding an
explicit choice is worse than refusing a combination that cannot work.

readme_threadx.txt carried the same claim, and its option list marked the FIQ
nesting dependency inline but not this one.  Both corrected.

No regression test.  Nothing in the tree asserts a configure-time failure --
there is no harness for it, and the sibling FIQ nesting guard has none either
-- so a test for this would have to introduce that mechanism for one case.
Verified by hand in both directions instead: the soft-ABI combination now
stops at configure with the message above, and the hard-ABI feature build
(VFP, FIQ, IRQ nesting, FIQ nesting) builds its eight images clean and passes
ctest 8/8 on the Armv8-R AEM FVP.  scripts/check_gcc.sh passes unchanged; it
configures the Cortex-R52 CMake stage without TX_R52_ENABLE_VFP, so the new
branch is not on its path.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-01 08:48:20 -04:00
Frédéric Desbiens a2800fef16 Covered the SMP suspension teardown and the long byte pool search (#677)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
Against the merged SMP coverage report -- every build configuration instrumented
and unioned -- sixty-four lines of common_smp/src were uncovered, 5114 of 5178.
Fifty-three of them are closed here and the report reads 5167 of 5178. The SMP
coverage floor goes from 98 to 99 with it.

Thirty-six of the sixty-four were one loop repeated four times: the walk in
tx_block_pool_delete, tx_byte_pool_delete, tx_event_flags_delete and
tx_queue_delete that releases every thread suspended on the object with
TX_DELETED. The suite deletes all four object types after every single test, and
that is exactly why the loop never ran. test_control_cleanup in the ThreadX
suite deletes the application's objects first and its threads last, so a test
that ends with a thread parked on a queue has that thread walked out of it by
tx_queue_delete. The SMP suite's cleanup deletes the threads first, deliberately
-- it was changed so that no application-owned object is still referenced when
the object loops run, which is what stopped a class of teardown hang. The side
effect is that all four deletes now run against an empty suspension list.
tx_semaphore_delete is the one member of the family that was already covered,
because threadx_semaphore_delete_test deletes a busy semaphore on purpose.

threadx_object_delete_suspension_test is the same idea for the other four. Two
threads suspend on each of a block pool, a byte pool, an event flags group and a
queue; the control thread waits on the object's own suspended count through
tx_*_info_get rather than on an ordering it cannot guarantee across four cores,
deletes the object, and checks both waiters came out with TX_DELETED. Two
waiters rather than one so the loop takes its back edge as well as its body, and
every wait is bounded in ticks so a suspension that never arrives fails the test
instead of hanging it.

threadx_trace_entry_update_test and threadx_thread_misaligned_stack_test are
ports of the two tests that closed the equivalent gaps in common/src, and close
fourteen more lines here: tx_block_allocate 123, 175, 182, 319 and 326,
tx_byte_allocate 130, 210, 217, 359 and 366, tx_thread_system_suspend 504 and
560, tx_trace_object_register 221, and tx_thread_create 133. The one substantive
change is core confinement. The trace test needs thread 0 to suspend and thread
1 to then release what it waits for; on four cores thread 1 gives the block back
before thread 0 has suspended and the update block behind the suspension is
never reached, so both threads are excluded from cores 1 to 3. The misaligned
stack test needed no such change.

threadx_byte_memory_long_search_test closes three of the eleven in
tx_byte_pool_search. Lines 264, 267 and 270 are the
TX_BYTE_POOL_MULTIPLE_BLOCK_SEARCH limit -- twenty on this port -- where a long
search drops and retakes protection so that it cannot lock the other cores out
for the whole walk. No byte pool in the suite ever had twenty fragments. This
one is filled with small chunks until it refuses another and then has every
second chunk released, so the free fragments are never adjacent and cannot be
merged, and the request is larger than any of them but smaller than the pool's
theoretical total, which is what makes _tx_byte_pool_search walk rather than
refuse at the door. The layout is asserted rather than assumed: the test checks
the fragment count and checks the probe request really does fail before the
workers start, because either would otherwise turn it into a silent no-op.

Eleven lines remain and they are not a to-do list. Eight are the delay loop in
tx_byte_pool_search that fires when another thread claims the pool inside the
window the search opens. The Linux SMP port serialises all four cores on one
pthread mutex, so that window is an unlock immediately followed by a lock on
that mutex, and glibc hands an uncontended mutex straight back to the thread
that just released it: measured over 180,003 windows across three cores, with
zero handovers. The shipped test therefore does 250 searches per worker rather
than the sixty thousand that probe used, because the twenty-block threshold is
crossed by the first search. The other three are in tx_thread_smp_utilities.
Line 149 is a range guard placed after the shift it is meant to guard, so
reaching it needs a shift by the width of the type; the fix is to move the check
above the shift, matching the TX_MAX_PRIORITIES > 32 variant of the same
function, and that belongs in its own change. Lines 1073 and 1074 need a mutex
owner that is genuinely executing on another core when a waiter suspends, and
three shapes were tried without producing one on this port.

Measured twice before and twice after, every gcda deleted between runs and
570 of 570 tests passing each time: 5114 of 5178 both times before, 5167 of 5178
both times after. Branch coverage goes from 2768 to 2821 and 2823 of 3548. A
floor of 99 needs 5127, so the ratchet lands with forty lines of headroom
against a numerator that has been seen moving by two between runs.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-28 10:41:39 -04:00
Frédéric Desbiens ff7fbe02f8 Added a GCC check for the Arm ports and ran it in CI (#675)
* Assembled the module ports, which no check had ever compiled

scripts/check_clang.sh globbed ports_module/*/gnu/src, which does not exist --
the module ports keep their assembly in module_manager/src. The [ -d ] guard
skipped it in silence, so 116 assembly files across nine Arm module ports were
assembled by no check, with either compiler, in the script whose own comments
state three times that "a port that is simply absent from the count reads as
covered". Stage 1 goes from 724 of 724 to 840 of 840; the feature-macro stage
had the same gap and goes from 412 files to 469.

Correcting the path exposed five defects, and only one of them was a build
failure. The other four assembled cleanly and did the wrong thing, because GAS
runs the C preprocessor on .S and not on .s:

  ports_smp/cortex_a7_smp/gnu/src/tx_thread_smp_unprotect.s, the only .s in a
  directory of twenty-one .S, ignored all four of its own feature macros. It
  wrote the caller's LR into the protection structure on every unprotect -- a
  store guarded by TX_MPCORE_DEBUG_ENABLE -- sent an unconditional SEV, and
  returned through both BX lr and MOV pc, lr. Its cortex_a5_smp and
  cortex_a9_smp siblings are .S.

  ports_module/cortex_m33/.../tx_thread_stack_build.s emitted both arms of an
  #ifdef TX_SINGLE_MODE_SECURE, so the non-secure LR value overwrote the secure
  one and the secure build got the wrong frame.

  ports_module/cortex_m23/.../tx_thread_context_{save,restore}.S carried the
  POP {r0, lr} that check_clang.sh's own comment describes as the reason the
  feature-macro stage exists. The 16-bit Thumb POP takes r0-r7 and pc only.
  The identical fix already sits in ports/cortex_m23/gnu/src; the module copy
  never got it because nothing scanned it.

  ports_module/cortex_m23/.../tx_thread_secure_stack_initialize.S used MOV
  rather than MOVS for an 8-bit immediate, latent behind TX_SINGLE_MODE_SECURE.
  Both siblings in the same directory already use MOVS.

  ports_module/cortex_a7/gnu/module_manager/src is the one that failed to
  assemble, on GCC 14.3 as well as on LLVM: #define SYS_MODE was never
  expanded, so #SYS_MODE reached the assembler as an undefined symbol.

Twenty-nine .s files under gnu trees are renamed to .S. Every one of them is
already named .S by the build scripts that compile it, so this repairs those
scripts rather than churning them -- ports_module/cortex_a7's build_threadx.bat
names all eighteen with a capital S, and works today only on a case-insensitive
filesystem. Renaming rather than converting the #defines to GNU assignments is
what fixes the #ifdef blocks as well as the constants; the assignments would
have fixed two files and left twenty-seven silently ignoring their macros.

Files with no preprocessor directives are left as .s: they are not broken, and
check_ports.sh gains a check that keeps them that way. Only the gnu trees are
checked there -- the IAR, Arm Compiler 5 and Keil assemblers preprocess .s
themselves, and about three hundred files in this repository rely on that.

Verified with both toolchains on the same tree: 840 of 840 assembled by
ATfE 22.1.0 and by arm-gnu-toolchain 14.3.rel1, all five stages of
check_clang.sh green, and check_ports.sh green including the reproducibility
check. The new check was shown to fail by planting a copy of the file it was
written for.

No regression test accompanies this. The assembly it covers is executed by no
host test, and the check itself going from 724 files to 840 is the coverage
AGENTS.md asks for -- together with the new check_ports.sh section, which is
what stops the class recurring.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Fixed the AArch64 samples, none of which had ever linked with GCC

Every AArch64 gnu example build failed at the sample link, all 27 of them --
13 under ports/ and 14 under ports_smp/:

  libg.a(libc_a-init.o): in function `__libc_init_array':
      undefined reference to `_init'
      relocation truncated to fit: R_AARCH64_CALL26 against undefined
          symbol `_init'
  libg.a(libc_a-fini.o): in function `__libc_fini_array':
      undefined reference to `_fini'

build_threadx_sample.sh links with -nostartfiles, which is correct for a port
carrying its own reset path, and that drops crti.o and crtn.o along with
everything else. startup.S calls __libc_init_array by design, and newlib's
implementation calls _init, which crti.o is what defines. The AArch32 scripts
are unaffected: they use nosys.specs and never reach __libc_init_array.

The fix links crti.o and crtn.o explicitly, bracketing the object list -- the
first must precede every .init contribution and the second must follow all of
them, so their position is load-bearing rather than stylistic. Both paths come
from the compiler's own -print-file-name, so nothing here hard-codes a
toolchain layout.

The atfe branch sets both to empty, deliberately: picolibc's __libc_init_array
does not call _init, those 27 images link today, and adding crti.o would change
a working link for no reason. That is also why check_clang.sh is green on these
and does not list them as expected to fail -- the LLVM path never reached the
gap, so nothing has ever linked them and failed.

Fixed in ports_arch/ARMv8-A/threadx/ports/gnu/example_build, which is the
single source for both the ports/ and ports_smp/ copies, then regenerated with
update.sh --port-sets tx,tx_smp. The 27 generated copies are in this commit
because ports_arch_check compares them.

Verified: all 27 link with arm-gnu-toolchain 14.3.rel1 aarch64-none-elf, where
0 of 27 did before; _init and _fini disassemble to the expected crti prologue
and crtn epilogue over a ret; check_clang.sh with ATfE 22.1.0 is still green on
all five stages, including the 42 script-driven example builds; check_ports.sh
is green including the reproducibility check.

No regression test: these are link-only example images that no host test
executes. What guards them is check_clang.sh's example stage today, and
check_gcc.sh's, which is the next change and is the reason this was found.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Added a GCC check for the Arm ports, which nothing had ever compiled

GCC is the project's declared default compiler (AGENTS.md, "The default
compiler for the project is GCC 14 on Linux"), it is what the gnu ports exist
for, and it is what nearly every downstream user builds with -- and nothing in
CI compiled a line of any port with it. The only cross-compilation check that
ran was the LLVM one, so the ATfE path was better guarded than the GNU one, on
ports whose directory is literally named gnu. ci_cortex_m covers four port
families; this covers forty.

Five stages, mirroring scripts/check_clang.sh stage for stage:

  1. assemble every .S and .s of every Arm gnu port -- 840 files
  2. assemble again the parts behind TX_ENABLE_VFP_SUPPORT,
     TX_ENABLE_FIQ_SUPPORT, TX_LOW_POWER and
     TX_ENABLE_EXECUTION_CHANGE_NOTIFY -- 469 files
  3. compile common/src for one core per architecture profile -- 185 x 9
  4. link the script-driven example builds -- 42
  5. link the CMake-driven Cortex-R52 images -- 5

Two scripts rather than one with a --toolchain flag: the flag surface differs
(a prefixed driver against --target=), the C library differs, and the set of
examples that can link differs. Folding them together makes it easy to weaken
one check while working on the other.

Two toolchains, both required. Arm ships arm-none-eabi and aarch64-none-elf as
separate downloads and PORT_TARGET maps every port to one of exactly those two
triples, so --arm-none-eabi and --aarch64-none-elf each take a driver or the
directory holding it, defaulting to the environment and then to PATH. A missing
one is a hard error rather than a soft skip: letting a run cover half the tree
and still report "all checks passed" is the failure this script exists to end.

PORT_TARGET is copied verbatim from check_clang.sh, including its warning not
to prefix-match core names -- cortex_a5* also matches the AArch64 cortex_a53.

VFP_EXTRA is the one map that is not a copy, and check_clang.sh's comment about
it is false for GCC. That comment says the A-profile defaults are already
correct; arm-none-eabi-gcc defaults to -mfloat-abi=soft, which disables the FPU
outright, so every VFP file fails with "selected processor does not support
'vmrs r1,FPSCR' in ARM mode". -mfloat-abi=hard alone is the fix and is the
right one, because it selects the core's own default FPU rather than naming a
-d16 one -- which is the trap the clang script warns about, since the
A-profile paths save D16-D31. Cortex-R4 is the exception in both scripts and
for the same reason: its FPU is an option rather than part of the core, so an
explicit -mfpu is required. Every value was measured against 14.3.rel1.

Stage 4 *unsets* TOOLCHAIN rather than setting it. The example build scripts
already default to GNU, and a stray TOOLCHAIN=atfe from a developer's shell
would otherwise make this stage silently check the other compiler. It cleans
the example directories on both sides, because the success test is the
existence of sample_threadx.out rather than the driver's exit status, and a
stale image from a previous toolchain would report success. Failure logs are
printed unfiltered: a missing tool says "command not found", and GNU ld's
undefined-symbol lines carry no "error:" at all.

Every skip is printed by name with a reason, per the house rule check_clang.sh
states three times -- a port simply absent from the count reads as covered.
This script also says outright that arm9 and arm11 are Arm and are skipped for
having no PORT_TARGET entry, which the clang script's "not Arm" wording glosses.

Verified on this tree with arm-gnu-toolchain 14.3.rel1: all five stages green,
every count identical to check_clang.sh's on the same tree -- 840, 469, 185x9,
42, 5 -- in 4m28s.

The failure paths were tested, not assumed. A deliberately broken .S in a
module port is reported by name and line in stages 1 and 2 and exits 1, in
--quiet mode as well. Reverting the AArch64 _init/_fini fix on one port only
gives "FAIL: cortex_a53: example build produced no image", 41 of 42, and exit
1 -- and the log tail it prints contains no "error:" anywhere, which is why it
is not filtered. A missing or wrong toolchain path exits 1 naming which triple
was not found.

RISC-V is deliberately out of scope for this first version: both ports
assemble 8 of 8 with the project's own cmake flags, but adding them widens the
toolchain download and the review surface for a family that is not regressing.

No regression test accompanies this. The script is the test, it exercises no
runtime behaviour, and its own failure paths are exercised above.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Ran the GCC port check in CI, on dev as well as master

scripts/check_gcc.sh with nothing invoking it would be a script nobody runs.
This adds the workflow, modelled on clang_check.yml, and fixes a trigger gap in
that file at the same time.

One job, two cache steps. Arm ships AArch32 and AArch64 as separate downloads
and the script needs both, so two caches keep the checks list short and let a
single invocation see both compilers. The AArch32 cache path and key match
cortex_m's exactly, so the two workflows share one entry rather than each
holding its own copy of the same archive -- noted in a comment, because the
only symptom of breaking that is a slower run.

Both triggers name dev. A workflow that triggers only on master gates no pull
request anybody opens; that is the defect ports_arch_check.yml carries a
comment about, and it cost cortex_m three months of failing in seven seconds
unnoticed. push is included as well as pull_request so dev's own history has a
baseline and a bad squash-merge is caught rather than waiting for the next PR.

The checksum suffix is .sha256asc and it is not interchangeable with .sha256.
Arm publishes both for this release, and verified 26 Aug 2026, the .sha256 file
for arm-none-eabi contains a 32-character MD5 rather than a SHA-256, so
sha256sum -c on it fails with "no properly formatted checksum lines found".
.sha256asc is a plain sha256sum-format line for both triples. The plan warned
that this suffix had changed between releases; the sharper truth is that both
suffixes exist simultaneously and one of them is not a SHA-256 at all. Recorded
in a comment beside the step.

Verified before writing them in rather than copied: both archive URLs and both
checksum URLs resolve, the archives are xz, the checksum files are
sha256sum-format for .sha256asc, and the AArch64 archive extracts to
arm-gnu-toolchain-14.3.rel1-x86_64-aarch64-none-elf/bin/aarch64-none-elf-gcc,
which is the path the workflow builds.

The paths: lists are duplicated between push and pull_request rather than
shared through a YAML anchor, deliberately: GitHub Actions' parser does not
dependably honour anchors and the failure mode is the workflow refusing to
parse, which is the cortex_m failure again. Ten duplicated lines are cheaper.

clang_check.yml's paths: list was missing CMakeLists.txt, cmake/ and
common_smp/, so that check did not run when files it reads changed -- the
ports_smp example builds compile common_smp/src and its CMake stage reads the
toolchain file and the top-level project. Both lists are now identical apart
from each file's own name, and both say so.

cortex_m is kept rather than deleted, against the plan's recommendation. It
builds four ports *through CMake*, and that is the only thing exercising
cmake/cortex_m*.cmake and the top-level CMakeLists for the M profile; this
script's CMake stage covers cortex_r52 only. The overlap is the assembly and
the C sources, not the build system, so deleting it would lose coverage rather
than remove a duplicate. Said so in the workflow header.

The script is passed explicit toolchain paths rather than left to find the
drivers on PATH, so nothing about the runner image can decide which compiler
runs, and it prints both versions it resolved.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-28 10:40:37 -04:00
Frédéric Desbiens 5b94bad6a2 Fixed the AArch64 samples, none of which had ever linked with GCC (#673)
Every AArch64 gnu example build failed at the sample link, all 27 of them --
13 under ports/ and 14 under ports_smp/:

  libg.a(libc_a-init.o): in function `__libc_init_array':
      undefined reference to `_init'
      relocation truncated to fit: R_AARCH64_CALL26 against undefined
          symbol `_init'
  libg.a(libc_a-fini.o): in function `__libc_fini_array':
      undefined reference to `_fini'

build_threadx_sample.sh links with -nostartfiles, which is correct for a port
carrying its own reset path, and that drops crti.o and crtn.o along with
everything else. startup.S calls __libc_init_array by design, and newlib's
implementation calls _init, which crti.o is what defines. The AArch32 scripts
are unaffected: they use nosys.specs and never reach __libc_init_array.

The fix links crti.o and crtn.o explicitly, bracketing the object list -- the
first must precede every .init contribution and the second must follow all of
them, so their position is load-bearing rather than stylistic. Both paths come
from the compiler's own -print-file-name, so nothing here hard-codes a
toolchain layout.

The atfe branch sets both to empty, deliberately: picolibc's __libc_init_array
does not call _init, those 27 images link today, and adding crti.o would change
a working link for no reason. That is also why check_clang.sh is green on these
and does not list them as expected to fail -- the LLVM path never reached the
gap, so nothing has ever linked them and failed.

Fixed in ports_arch/ARMv8-A/threadx/ports/gnu/example_build, which is the
single source for both the ports/ and ports_smp/ copies, then regenerated with
update.sh --port-sets tx,tx_smp. The 27 generated copies are in this commit
because ports_arch_check compares them.

Verified: all 27 link with arm-gnu-toolchain 14.3.rel1 aarch64-none-elf, where
0 of 27 did before; _init and _fini disassemble to the expected crti prologue
and crtn epilogue over a ret; check_clang.sh with ATfE 22.1.0 is still green on
all five stages, including the 42 script-driven example builds; check_ports.sh
is green including the reproducibility check.

No regression test: these are link-only example images that no host test
executes. What guards them is check_clang.sh's example stage today, and
check_gcc.sh's, which is the next change and is the reason this was found.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-28 09:26:12 -04:00
Frédéric Desbiens 9c32abb17d Assembled the module ports, which no check had ever compiled (#672)
scripts/check_clang.sh globbed ports_module/*/gnu/src, which does not exist --
the module ports keep their assembly in module_manager/src. The [ -d ] guard
skipped it in silence, so 116 assembly files across nine Arm module ports were
assembled by no check, with either compiler, in the script whose own comments
state three times that "a port that is simply absent from the count reads as
covered". Stage 1 goes from 724 of 724 to 840 of 840; the feature-macro stage
had the same gap and goes from 412 files to 469.

Correcting the path exposed five defects, and only one of them was a build
failure. The other four assembled cleanly and did the wrong thing, because GAS
runs the C preprocessor on .S and not on .s:

  ports_smp/cortex_a7_smp/gnu/src/tx_thread_smp_unprotect.s, the only .s in a
  directory of twenty-one .S, ignored all four of its own feature macros. It
  wrote the caller's LR into the protection structure on every unprotect -- a
  store guarded by TX_MPCORE_DEBUG_ENABLE -- sent an unconditional SEV, and
  returned through both BX lr and MOV pc, lr. Its cortex_a5_smp and
  cortex_a9_smp siblings are .S.

  ports_module/cortex_m33/.../tx_thread_stack_build.s emitted both arms of an
  #ifdef TX_SINGLE_MODE_SECURE, so the non-secure LR value overwrote the secure
  one and the secure build got the wrong frame.

  ports_module/cortex_m23/.../tx_thread_context_{save,restore}.S carried the
  POP {r0, lr} that check_clang.sh's own comment describes as the reason the
  feature-macro stage exists. The 16-bit Thumb POP takes r0-r7 and pc only.
  The identical fix already sits in ports/cortex_m23/gnu/src; the module copy
  never got it because nothing scanned it.

  ports_module/cortex_m23/.../tx_thread_secure_stack_initialize.S used MOV
  rather than MOVS for an 8-bit immediate, latent behind TX_SINGLE_MODE_SECURE.
  Both siblings in the same directory already use MOVS.

  ports_module/cortex_a7/gnu/module_manager/src is the one that failed to
  assemble, on GCC 14.3 as well as on LLVM: #define SYS_MODE was never
  expanded, so #SYS_MODE reached the assembler as an undefined symbol.

Twenty-nine .s files under gnu trees are renamed to .S. Every one of them is
already named .S by the build scripts that compile it, so this repairs those
scripts rather than churning them -- ports_module/cortex_a7's build_threadx.bat
names all eighteen with a capital S, and works today only on a case-insensitive
filesystem. Renaming rather than converting the #defines to GNU assignments is
what fixes the #ifdef blocks as well as the constants; the assignments would
have fixed two files and left twenty-seven silently ignoring their macros.

Files with no preprocessor directives are left as .s: they are not broken, and
check_ports.sh gains a check that keeps them that way. Only the gnu trees are
checked there -- the IAR, Arm Compiler 5 and Keil assemblers preprocess .s
themselves, and about three hundred files in this repository rely on that.

Verified with both toolchains on the same tree: 840 of 840 assembled by
ATfE 22.1.0 and by arm-gnu-toolchain 14.3.rel1, all five stages of
check_clang.sh green, and check_ports.sh green including the reproducibility
check. The new check was shown to fail by planting a copy of the file it was
written for.

No regression test accompanies this. The assembly it covers is executed by no
host test, and the check itself going from 724 files to 840 is the coverage
AGENTS.md asks for -- together with the new check_ports.sh section, which is
what stops the class recurring.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-28 09:24:30 -04:00
Frédéric Desbiens 147754cc86 Enforced a coverage floor on the merged report (#667)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The coverage summary reported a percentage and could not fail. Coverage
could fall from 99.97% to anything at all and every check stayed green,
against an AGENTS.md that asks for 100% test coverage -- a stated
requirement measured with a gauge that had no failure mode.

CodeCoverageSummary already takes thresholds and fail_below_min; neither
was set. Both are now, through a new coverage_thresholds input on the
template, because the two suites do not sit at the same figure: ThreadX
99, SMP 98.

Three things were probed against the pinned action on a runner before
picking those numbers, using the real merged.xml files from the dev push
run of #666.

The floor compares the line rate and nothing else. That mattered because
branch coverage is around 78% in both suites while line coverage is
98.8-100%, so a floor aimed at the line figure would have been an
immediate red wall had it tested branches or the lower of the two. The
ThreadX report at 100.00% lines and 77.67% branches clears a floor of 99.

The thresholds are whole numbers. '99.9 100' -- the value this was meant
to be -- is rejected with 'System.ArgumentException - Threshold parameter
set incorrectly.', and the step fails whether or not fail_below_min is
set. So the choice is 99 or 100 with nothing between.

100 would fail on a race. tx_thread_system_resume.c:529 is reached by
timing rather than by construction and flaps between runs of the same
green tree, which is why #666 left it; 4502/4503 fails a floor of 100 and
clears one of 99. A coverage gate that goes red on a coin toss is how
coverage gates get switched off.

SMP is 5114/5178 lines, 98.76%, with 64 uncovered lines across 11 files
of common_smp/src -- #666 closed the equivalent gaps in common/src only.
A shared floor of 99 would have failed that job on every run while
ThreadX passed.

One limit is recorded in the file rather than fixed: an empty report
reads as 100%. gcovr writes line-rate="1.0" beside lines-valid="0" when
it finds no data, and the action prints 'Line Rate = 100% (0 / 0)' and
passes any floor. The check for that is the emptiness assertion #664 put
in each suite's coverage.sh, not this one.

Also corrected two stale filenames in the deploy job's comment: since
#665 each coverage artifact carries merged.xml, not
default_build_coverage.xml. Verified on the runner.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 09:13:35 -04:00
Frédéric Desbiens f89d65f041 Covered the trace entry update paths and the misaligned stack adjustment (#666)
Against the merged coverage report -- every build configuration instrumented and
unioned -- eighteen lines of common/src were uncovered. Seventeen of them were
one missing scenario rather than eighteen separate gaps.

tx_block_allocate, tx_byte_allocate, tx_thread_system_suspend and
tx_thread_system_resume each carry blocks under TX_ENABLE_EVENT_TRACE that go
back and patch a trace entry once the call has done its work, all of the shape:

    if (entry_ptr != TX_NULL)
    {
        if (time_stamp == entry_ptr -> tx_trace_buffer_entry_time_stamp)

entry_ptr comes from _tx_trace_buffer_current_ptr, which stays TX_NULL until
tx_trace_enable is called at run time. Building with TX_ENABLE_EVENT_TRACE is
not enough, and exactly one test in the suite enables tracing --
threadx_trace_basic_test -- which tests the enable API itself and never calls
either allocator. So those blocks sat in the report's denominator and never in
its covered set.

threadx_trace_entry_update_test enables tracing and then drives both allocators
twice each, once on the path that succeeds immediately and once through a
suspension that a second thread satisfies, since each allocator carries one
update block on either side. It then sleeps so that the last runnable thread
suspends with nothing ready to take over: tx_thread_system_suspend lines 345 and
351 are on the branch that sets _tx_thread_execute_ptr to TX_NULL, and the two
allocator suspensions never reach it because the other thread was always ready.

The same test closes tx_trace_object_register's TX_NULL name break by creating a
semaphore with no name. A semaphore and not a thread deliberately: for
TX_TRACE_OBJECT_TYPE_THREAD the register function dereferences the pointer it is
given to read the thread's priority, so that type needs a real TX_THREAD behind
it. threadx_trace_basic_test makes the equivalent call only under
ifndef TX_ENABLE_EVENT_TRACE, against the no-op stub.

threadx_thread_misaligned_stack_test covers the remaining line,
tx_thread_create.c:136, where a stack that does not begin on a ULONG boundary
costs a ULONG of size so that rounding the start up cannot run past the end of
the caller's buffer. Every other test hands tx_thread_create an aligned stack.
That line is compiled only under TX_ENABLE_STACK_CHECKING, so it is absent from
three of the five configurations' reports rather than uncovered in them, and it
was verified under stack_checking_build.

Measured on the merged report, all 480 tests passing and 5 of 5 configurations
green: 4485 of 4503 covered before, 4502 of 4503 after, denominator unchanged.

One line remains, tx_thread_system_resume.c:529, and it is the report's last
flapping line rather than a standing gap -- two clean runs of the same tree gave
4503 of 4503 and 4502 of 4503. Reaching it by construction was tried twice and
failed both times, so it is left alone here. It needs the preempt disable flag
and the system state both clear, and tx_thread_resume raises the preempt disable
flag before calling _tx_thread_system_resume, as do the put and send paths;
creating a higher priority auto-start thread from thread context does not raise
it but does not reach the check either, which a probe showed is executed only
during initialisation, with the system state at TX_INITIALIZE_IN_PROGRESS.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 08:14:38 -04:00
Frédéric Desbiens 3e85bbd431 Instrumented every build configuration and merged their coverage (#665)
Only default_build_coverage carried -fprofile-arcs, because the gate was the
build type and it is the only one of five whose name ends in _coverage. The
other four build and run all their tests and their coverage was discarded. That
is not redundancy thrown away: each configuration selects a different set of TX_
feature macros, so the code the other four compile is absent from the
denominator rather than uncovered in it.

TX_COVERAGE instruments a build regardless of its name, defaulting to OFF so a
single configuration built by hand behaves as before. coverage.sh gains a
--merge mode that unions the per-configuration JSON tracefiles, and
cmake_bootstrap.sh runs it after the test loop so a local run produces the same
merged report CI reads. The template sets TX_COVERAGE for build and test, and
coverage_name moves to the merged report.

Measured on the ThreadX suite, all 480 tests passing:

  default_build_coverage           3827 valid   3827 covered
  disable_notify_callbacks_build   3767         3766
  stack_checking_build             3857         3856
  stack_checking_rand_fill_build   3862         3861
  trace_build                      4123         4108
  merged                           4503         4487

The denominator grows by 676 lines, 17.7%, and the figure moves from 99.97% to
99.64%. The second one is honest, and the drop is the point rather than a
regression: the denominator now includes code the old report never counted. The
union also contains a file the old report did not contain at all --
tx_thread_stack_error_handler.c compiles only under TX_ENABLE_STACK_CHECKING, so
it was not listed at 0%, it was simply absent. 177 files becomes 178.

Coverage collection moved out of test() and now runs after the test loop, one
configuration at a time. gcov writes its intermediate gcov files into the
directory gcovr is rooted at, and coverage.sh roots every configuration at the
repository root so filenames come out repo-relative. Five concurrent gcovr
processes therefore share one scratch directory and delete each other's output:
the first full run of this change passed all 480 tests and produced no report
for three of the five configurations. Measured both ways -- two gcovr rooted at
the repository root fail concurrently and succeed in sequence. CI would not
have caught it, because test_tx.sh sets CTEST_PARALLEL_LEVEL=1 and takes the
serial branch.

Per-configuration output moved under coverage_report/per_configuration/ and is
excluded from the Pages artifact. The deploy job merges the ThreadX and SMP
artifacts into one tree and every configuration directory has the same name in
both, so left at the top level one suite's would overwrite the other's on the
published site.

On the SMP suite, an earlier run of this change saw trace_build fail
threadx_smp_time_slice_test and then hang, which raised the question of whether
-fprofile-arcs perturbs a timing-sensitive test. It does not. Sixteen runs
settle it, and the decisive one is that threadx_smp_time_slice_test failed
ERROR #31 -- twice in a row under --repeat until-pass:2 -- on an uninstrumented
build, in the exact shape CI runs, while three instrumented runs of that shape
passed 5 of 5. In the CI shape, CTEST_PARALLEL_LEVEL=1 run.sh test all:

  TX_COVERAGE=OFF   3 runs   2 green, one ERROR #31        310 s
  TX_COVERAGE=ON    3 runs   3 green, 5/5 each             325-329 s

So the test is a pre-existing flake on dev and instrumenting all five costs
about 5% of the suite's wall clock. Separately, and also in both instrumented
and uninstrumented builds, run.sh's parallel branch -- what a developer gets
typing run.sh test all with no CTEST_PARALLEL_LEVEL -- hangs under its own load,
four times in twelve runs. Several SMP tests create 1024 ThreadX threads by
construction and the Linux port backs each with a pthread, so five
configurations at once put on the order of 5000 threads on the machine. CI sets
CTEST_PARALLEL_LEVEL=1 and does not take that branch.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 08:13:33 -04:00
Frédéric Desbiens b756220c43 Fixed the coverage report's paths and scoping (#664)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The Cobertura XML embedded absolute machine paths, and the flag that looked
like it scoped the report to one build configuration was doing nothing at all.
Both coverage.sh scripts had the defect; both are fixed here, because the SMP
report is published to the same Pages site as the ThreadX one.

Paths. -r was the build directory and -f pointed outside it, so gcovr could not
express the sources relative to the root and fell back to absolute paths. The
result named files as /home/runner/work/threadx/threadx/common/src/... while
the <source> element beside them said build/default_build_coverage, so the two
halves of the same file disagreed and nothing could map coverage back to the
repository. -r is the repository root now and -f an absolute path beneath it,
which gives filename="common/src/tx_block_allocate.c".

Both must be absolute: -r ../../.. -f common/src produces a report of zero
files and exits 0, which is the worst failure mode available here.

Scoping. --object-directory does not restrict which gcda files are found -- it
tells gcovr how to get from a gcda file back to the compiler's working
directory. Pointed at an empty directory it still produced the full 177-file
report. That was harmless only by accident, because -r build/$1 constrained the
search instead; moving -r to the repository root removes that accident, so the
two changes have to land together. Measured, with a second instrumented
configuration deliberately made sparser than the first:

  scoped by the positional search path   3221 of 3827 lines -- the truth
  no search path, -r at the repo root    3827 of 3827 -- silently merged
  --object-directory at the sparse tree  3827 of 3827 -- scopes nothing

So the search path is load-bearing, and it matters ahead of instrumenting all
five configurations: without it each configuration would have reported the
union as its own.

An empty report is not an error to gcovr -- it warns and exits 0 -- and it
carries line-rate="1.0" next to lines-valid="0", so a consumer reads no data at
all as fully covered. No coverage threshold can catch that, since an empty
report passes any threshold. Hence the explicit assertion that the report has
content, which fires with exit 1 on an object directory that exists but is
empty, where the old shape returned 177 files and exit 0.

Also says out loud that ports/linux/gnu/src is deliberately outside the filter.
gcno files exist for it and it is dropped without a word today.

Number-neutral, and that was the test. Over the same frozen gcda, changing only
the gcovr invocation: ThreadX 3827 of 3827 lines and 1993 of 1994 branches
across 177 files, SMP 4739 of 4791 and 2417 of 2430 across 185, before and
after alike. Same answers on gcovr 7.0, 8.3 and 8.6, so the change is not
wedged to the current pin. End to end through run.sh, 96 of 96 and 110 of 110
pass with the reports written.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 17:16:51 -04:00
Frédéric Desbiens 2d9b9f7417 Bumped gcovr off the 4.1 pin it had been held on since 2018 (#663)
The coverage tooling was pinned to gcovr 4.1, released in 2018, and that
version is missing the two options the coverage work needs next: --json and
--add-tracefile, which is how the five build configurations get merged into
one report. This moves the pin to 8.6, the current release. The pin stays
exact, and it stays hand-moved: it lives in a shell script, and no Dependabot
ecosystem can parse that.

Isolated deliberately, so that a movement in the coverage number caused by the
tool could not be confused with one caused by a later change. Measured on the
default_build_coverage tree of test/tx, over the same gcda with the same gcov,
varying only the gcovr version:

  gcovr    lines-valid  branches-valid  files
  4.1      3827         1994            177
  7.0      3827         1994            177
  8.3      3827         1994            177
  8.6      3827         1994            177

So the denominator does not move with the tool at all, and this bump moves no
number. The plan this came from expected 3822 to become 3827; that figure does
not reproduce, under gcc-13 or gcc-14, with or without --object-directory. The
only variant that changes the count is dropping the -f filter, which collapses
the report to nothing.

Two things found while measuring, both recorded because they matter to what
comes next.

The coverage numerator is not deterministic. On an identical tree with an
identical compiler, three consecutive runs of the full suite -- all 96 tests
passing every time -- reported 3826, 3827 and 3827 covered lines. The line that
flickers is tx_thread_system_resume.c:529, the preemption path of
_tx_thread_system_resume, and it takes its guarding branch with it. It has been
described as never executed; it is executed on some runs and not others. A
coverage floor has to be set with that in mind, and the honest fix is a test
that takes the path deliberately.

Reading gcc-13 output, the compiler the runners actually use, gcovr 8.6 runs
the existing coverage.sh unchanged: Cobertura XML and 181 HTML files, same 177
classes. --xml-pretty and --object-directory still work on 8.6 but are now
deprecated aliases for --cobertura-pretty and --gcov-object-directory, worth
knowing for whoever removes --object-directory next.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 16:35:16 -04:00
Frédéric Desbiens b6a00a2014 Added the Dependabot configuration the pinned actions need (#662)
The action references were pinned to commit SHAs in #660, and a SHA pin with
nothing moving it is worse than a floating tag -- it holds CI on whatever was
current the day it was written. That is exactly how actions/cache@v1 stayed in
ci_cortex_m.yml until GitHub began auto-failing every request that used it.
The drift measured before that catch-up: download-artifact four majors behind,
checkout and upload-artifact three each, cache and upload-pages-artifact two,
with nothing ever reporting it. This closes the loop, and the reference to
.github/dependabot.yml that #660 left in each workflow's pinning comment.

Weekly, github-actions only. Patch and minor are grouped into one pull request
because they are the routine traffic and a queue reviewed one item at a time is
a queue that gets ignored. Majors stay ungrouped, one each, because every
breaking change this repository has met in an action has been a major.

Two choices worth stating rather than leaving to be rediscovered.

target-branch is dev. Dependabot reads this file from the default branch, which
is master, but master is deliberately kept behind dev and pull requests belong
where the regression suites gate them. The consequence is that landing this on
dev arms it without firing it: nothing happens until a release merge carries
the file to master. Setting target-branch also opts out of Dependabot security
updates, which only run against the default branch -- a small cost for this
ecosystem, since an action advisory arrives as an ordinary bump on the weekly
run, but a real one.

The pull-request limit is raised from the default five to ten. Nine actions are
in use, and five would hold majors back with nothing saying that it had.

No other ecosystem is configured, deliberately: external dependencies are
forbidden, there are no submodules, and the one pinned tool -- gcovr in
scripts/install.sh -- lives in a shell script no ecosystem can parse, so that
pin keeps moving by hand.

No sibling eclipse-threadx repository has a Dependabot configuration, so this
sets the pattern rather than following one. The dependencies label it uses
already exists here.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 16:05:04 -04:00
Frédéric Desbiens 3d852eb451 Pinned every action to a commit SHA, and moved them off Node 20 (#660)
Node 20 is removed from the GitHub runners on 16 September 2026. Every run
in this repository currently emits the deprecation warning for it, naming
actions/checkout, actions/configure-pages, actions/upload-artifact,
LouisBrunner/checks-action and marocchino/sticky-pull-request-comment among
others. After that date those actions stop working rather than warning, so
this is a deadline and not housekeeping.

Every action is now referenced by a 40-character commit SHA with the version
in a trailing comment. A tag can be repointed at any commit; a SHA cannot, so
this is what makes "which code ran in our CI" answerable from the repository
rather than from whatever the tag meant at the time. The versions were behind
by as much as four majors -- download-artifact was on v4.3.0 against v8.0.1 --
because nothing in this repository has ever reported that an action moved.

Compatibility was checked against each new action.yml rather than assumed,
for every input this repository actually passes:

  checkout            submodules is unchanged
  cache               path and key are unchanged
  upload-artifact     name, path and retention-days are unchanged
  download-artifact   pattern, merge-multiple and path are unchanged
  configure-pages     takes no input here, and none became required
  deploy-pages        still exposes page_url, which the job reads
  upload-pages-art.   path is unchanged
  checks-action       token, name, conclusion, output and
                      output_text_description_file all survive v2 to v3
  sticky-comment      header and path survive v2 to v3, and the new
                      GITHUB_TOKEN input defaults to github.token, which is
                      what v2 used implicitly
  delete-artifact     name survives v5 to v6, and useGlob still defaults to
                      true, so the coverage_report-* glob from #655 still
                      matches
  CodeCoverageSummary already current at v1.3.0; pinned, not moved

The artifact pair moves together, as it must. The round trip was verified on
a runner before this commit: upload-artifact v7 to download-artifact v8,
through the pattern and merge-multiple selection #655 introduced, filtered 4
artifacts to 2 and produced exactly the tree the deploy expects.

Two behaviour changes worth knowing. download-artifact v8 adds a
digest-mismatch input defaulting to error, so a corrupted artifact now fails
the job instead of passing through -- the right default, but a change.
upload-artifact v6 and above require a runner of at least 2.327.1, which the
hosted runners satisfy and a self-hosted runner would need checking for.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 15:52:49 -04:00
Frédéric Desbiens 042049a7b2 Revived the Cortex-M build, which had compiled nothing since June (#653)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
Two defects, and the second hid the first.

This workflow triggered on master only, for both push and pull_request,
while dev is the integration branch. So it gated no pull request that
anybody opened. That is the same defect ports_arch_check.yml carries a
comment about, where it cost eight months of ports drifting from ports_arch
unnoticed, and regression_test.yml has it too.

And it had not compiled anything since at least 2026-06-08. Every run since
then failed in six to eight seconds at "Prepare all required actions",
before checkout, because GitHub automatically fails any request that uses
actions/cache@v1. The last run of any kind was 2026-06-30. A workflow that
fails in seven seconds is normally noticed within the day; this one was not,
because of the first defect. The two together meant the project's only job
that cross-compiles a port with GCC had been reporting nothing at all.

The toolchain now follows clang_check.yml rather than third party actions:
a pinned release fetched directly from Arm, verified against the published
sha256asc, and cached with actions/cache@v4. That also moves the compiler
off 9-2019-q4, a 2019 release, onto a version matching the GCC 14 default
this project states.

The ninja install is guarded on ninja being absent rather than run
unconditionally, because scripts/install.sh already carries a long comment
about apt-get update stalling for over two hours and taking a whole
regression run with it.

fail-fast is off so that one port failing still reports the other three.

Verified before committing, with the pinned 14.3.rel1 toolchain: the
download and sha256sum -c sequence in the install step was run as written,
the archive extracts to the directory the PATH step expects, and all four
ports configure and build clean with zero warnings.

This covers four ports of the forty under ports/ that have a gnu directory.
Widening it to every Arm gnu port is separate work.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 16:38:24 -04:00
Frédéric Desbiens 7959aef3bc Kept the coverage report from the runs that most need one (#659)
A failing test threw away coverage that had already been collected, and the
run whose behaviour changed is exactly the run whose coverage is worth
reading. Measured on the failing run of 2026-08-18: it uploaded test_reports
for all three suites and no coverage_report artifact at all.

Two causes, and the workflow one is the smaller of them.

cmake_bootstrap.sh runs under set -e, so a failing ctest aborted test()
before ./coverage.sh was reached. The gcda files exist by that point, so
nothing was missing except the step that reads them. ctest's status is now
captured and returned at the end, and the summary grep is allowed to fail
rather than being the thing that stops the coverage behind it.

The serial branch of the test dispatch collected no status either, so under
set -e the first failing configuration stopped the remaining four from being
tested at all -- and their coverage from being collected. That was cheap
while the suites ran in parallel, because the parallel branch already
collects exit codes from its background jobs. Moving to serial execution in
#643 quietly made one failure cost the other four configurations. The serial
branch now collects status the same way the parallel branch does.

With those fixed the report exists, so the workflow steps that publish it no
longer skip on failure. They are guarded with !cancelled() rather than
always(), so a cancelled run still stops promptly, which is the idiom
deploy_code_coverage already uses. The ${{ }} wrapping is required and not
decoration: a bare ! opens a YAML tag, and the file will not parse without
it.

Verified locally by replacing one test binary with a stub that exits 1:

  before  the failing run of 2026-08-18 produced no coverage_report artifact
  after   run.sh test default_build_coverage exits 8, and produces
          coverage_report/default_build_coverage.xml with 177 files and
          3804 of 3827 lines
  after   run.sh test all exits 8, and all five configurations run rather
          than stopping at the first

The failure still fails. Only the reporting around it changed.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 16:36:18 -04:00
Frédéric Desbiens adc6469b91 Published the coverage report instead of everything the run produced (#655)
The download step in deploy_code_coverage asked for the artifact named
${{ steps.artifact.outputs.coverage_report }}. That output is set by the
"Coverage Report name" step of run_tests, which is a different job, and the
steps context does not cross jobs. So the expression evaluated to the empty
string and the action took its documented path for an unspecified name:

    No input name, artifact-ids or pattern filtered specified,
    downloading all artifacts
    Total of 4 artifact(s) downloaded

The four are the two coverage reports and the two test_reports bundles of
JUnit XML, each extracted into a directory named after the artifact. The
next step uploads the lot to Pages, so the published site has carried the
test reports alongside the coverage, one directory deeper than intended,
under a path containing a run timestamp that changed on every publish. Any
link to a coverage report broke the next time one was published.

Selecting by pattern with merge-multiple fixes both halves: the pattern
excludes the test_reports bundles, and merging puts the contents of the two
coverage artifacts directly into coverage_report rather than under a
directory named for each. Each artifact holds one directory named for its
suite, renamed from default_build_coverage by "Prepare Coverage GitHub
Pages", so the result is the two suite directories the deploy expects and
the timestamped artifact name no longer appears in the published path.

Verified on a runner rather than reasoned about, with an isolated workflow
that uploads artifacts shaped like the real ones and downloads them both
ways:

    OLD  coverage_report/coverage_report-<epoch>-ThreadX/ThreadX/index.html
         coverage_report/coverage_report-<epoch>-ThreadX/default_build_coverage.xml
         coverage_report/coverage_report-<epoch>-SMP/SMP/index.html
         coverage_report/coverage_report-<epoch>-SMP/default_build_coverage.xml
         coverage_report/test_reports SMP/results.xml
         coverage_report/test_reports ThreadX/results.xml

    NEW  coverage_report/ThreadX/index.html
         coverage_report/SMP/index.html
         coverage_report/default_build_coverage.xml

Both artifacts carry a default_build_coverage.xml and the merge means one
overwrites the other, which the run above also shows. That file is consumed
by CodeCoverageSummary back in run_tests and is not read here, so it is
untidy rather than wrong, and it is called out in a comment.

The delete step is fixed in the same place and for a related reason. The
artifacts are named coverage_report-<epoch>, useGlob defaults to true in
this action, and as a glob "coverage_report" matches only the literal
string. It has been deleting nothing, without failing, and retention-days: 1
on the upload is what has actually been clearing these up.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 16:31:21 -04:00
Frédéric Desbiens eabdb86409 Matched gcov to the compiler that produced the coverage data (#658)
gcov reads a data format tied to the compiler that produced it. coverage.sh
took whatever gcov was first on PATH, which was fine while the compiler was
also whatever was first on PATH. #656 made cmake/linux.cmake honour CC and
#657 made a compiler switch actually reconfigure the build, so that
assumption no longer holds, and the first person to use the new capability
would have hit this.

Measured on dev with both of those merged:

    CC=gcc-14 ./run.sh build default_build_coverage    # succeeds
    ./coverage.sh default_build_coverage               # exit 64

gcov says why, if asked directly:

    tx_block_allocate.c.gcno:version 'B42*', prefer 'B33*'

gcovr turns that into "GCOV returncode was 3" and exits 64 through a Python
traceback, after the tests have already passed. It reads like a coverage bug
rather than a toolchain mismatch, which is the part that would have cost
someone an afternoon.

gcov is now derived from CC rather than found on PATH, so the caller sets
one variable instead of remembering two. GCOV still overrides, for a
toolchain that does not follow the gcc/gcov naming, and a derived gcov that
does not exist is reported as such instead of surfacing as a traceback.

Verified, tx and smp, before and after:

    CC=gcc-14    was exit 64, now exit 0, 177 files and 1527/3827 lines
    CC unset     exit 0, 177 files and 1527/3827 lines, unchanged
    CC=gcc-99    exit 1 naming gcov-99 and CC, rather than a traceback
    GCOV=gcov-14 with CC=gcc-99, exit 0, so the override still wins

A mismatched pairing still fails, deliberately: reading a gcc-14 tree with
the default gcc-13 gcov is exit 64 as before. Producing a number from
mismatched data would be worse than refusing.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 15:55:18 -04:00
Frédéric Desbiens a4397e133a Reconfigured the build when the requested compiler changes (#657)
CMake records the compiler it detected inside the build directory and keeps
using it on every later configure. Since cmake/linux.cmake began honouring CC,
that made a compiler switch silently ineffective: CC=gcc-14 ./run.sh build <cfg>
against an existing build directory printed "ninja: no work to do", exited 0,
and left the previous compiler in place. Anyone verifying a change against a
second compiler would have been reading stale results while being told the
build had succeeded.

generate() now compares the compiler recorded in the build directory with the
one currently requested, and reconfigures from scratch when they differ. The
comparison uses the path as CMake records it, which is the unresolved path as
given, so /usr/bin/gcc matches command -v gcc rather than the versioned target
its symlink points at. build_libs() gets the same treatment.

Only the C compiler is consulted: both trees that use this script declare
LANGUAGES C, so no CXX compiler is ever detected.

Nothing is reconfigured unless the compiler actually changed, so repeat builds
stay incremental and the default path is unchanged.

Assisted-by: Claude Code (Opus 5)
2026-08-24 14:47:26 -04:00
Frédéric Desbiens 5a68c9da4b Allowed the Linux toolchain file to accept a compiler override (#656)
cmake/linux.cmake set CMAKE_C_COMPILER and CMAKE_CXX_COMPILER unconditionally.
CMake reads a toolchain file before it consults CC and CXX, and a plain set()
in a toolchain file also takes precedence over -DCMAKE_C_COMPILER, so neither
of the two usual ways to pick a compiler had any effect: the tree could only
ever be built with whatever /usr/bin/gcc happened to point at.

That matters because AGENTS.md names GCC 14 as the project's default compiler
on Linux, while distributions still ship an older gcc as the default for some
time. Selecting GCC 14 previously meant either editing this file or changing
the machine's system-wide default.

Both variables now fall back to gcc and g++ only when nothing else has been
specified, so the default build is byte-for-byte what it was, while
-DCMAKE_C_COMPILER=gcc-14 or CC=gcc-14 now work as expected.

The binutils variables in this file are left alone: they are unused on the
Linux target, so guarding them would be unrelated churn.

Assisted-by: Claude Code (Opus 5)
2026-08-24 13:42:31 -04:00
Frédéric Desbiens 9218bad4bc Kept the coverage publish on master, where the environment allows it (#654)
Running the regression suites on dev (#652) was meant to test the branch the
pull requests target. It changed what gets published as well, which was not
intended and does not work: the first push to dev after that merge failed
with

    Branch "dev" is not allowed to deploy to github-pages due to
    environment protection rules.

All three suites passed in that run -- tx, smp and freertos. The only
failure was deploy / deploy_code_coverage, rejected before it ran, because
the github-pages environment restricts deployments to master.

The guard goes here rather than in the environment settings, because the
environment rule is doing its job. Which branch the published coverage
report describes is a deliberate decision, and moving it from master to dev
is a change worth making on purpose rather than as a side effect of a
trigger fix. Doing so needs the environment setting relaxed as well as this
line removed.

The per-suite deploy_code_coverage jobs need no guard: tx, smp and freertos
all pass skip_deploy: true, and regression_template.yml already restricts
that job to push and workflow_dispatch. Only the deploy job, which is the
one that publishes, was reaching the environment.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 13:07:38 -04:00
Frédéric Desbiens 977e14e776 Ran the regression suites on dev, where the pull requests actually are (#652)
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The ThreadX, SMP and FreeRTOS-compatibility suites triggered on master only,
for both push and pull_request. dev is the integration branch, so these
suites gated no pull request that anybody opened: the last dev run of any
kind was a manual workflow_dispatch on 2026-08-18.

This is the same defect ports_arch_check.yml already carries a comment
about, where it cost eight months of ports drifting from ports_arch
unnoticed. ci_cortex_m.yml has it too and is handled separately.

That 2026-08-18 run failed, which is the reason to check before switching
this on rather than after. Two tests failed: threadx_timer_simple_test in
the ThreadX suite, with ERROR #28, and threadx_thread_priority_change in
the SMP suite, with a timeout. Both were fixed two days later -- the first
by running the suites one test at a time (#643), which is what a timer test
failing only under parallel load wants, and the second by #647 by name.

Verified before this commit rather than assumed: both suites were re-run on
this tree, and all 1030 tests pass across all ten build configurations, the
ThreadX suite in 34 to 64 seconds per configuration and the SMP suite in 61
to 63. The 2026-08-18 run took 36m19s, of which a single test that has since
been given a budget accounted for 439 seconds.

No paths filter is added deliberately. The suites build the linux port, so a
filter would have to enumerate what cannot affect them, and the failure mode
of getting that list wrong is a regression that merges because the filter
excluded the file that caused it.

The deploy job needs no guard: regression_template.yml already restricts
deploy_code_coverage to push and workflow_dispatch, and restricts the
coverage PR comment to pull requests from the repository itself, so neither
fires for a pull request from a fork.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 11:30:19 -04:00
Frédéric DesbiensandClaude Opus 5 e46b1b0787 Asked the wait abort test for three windows, and failed a run that reached none (#649)
Four CI runs of the same tree, twenty configuration-runs in total, show this
test's budget being reached far more often than the first green run suggested,
and a pass being reported every time it was:

    trace_build            3 of 10 windows in 121 seconds
    disable_notify         3 of 10 windows in 121 seconds
    default_coverage       4 of 10 windows in 121 seconds
    stack_checking         7 of 10 windows in 121 seconds
    trace_build            0 of 10 windows in 121 seconds
    disable_notify         7 of 10 windows in 121 seconds
    stack_checking         3 of 10 windows in 121 seconds

Seven of twenty, and the shortfall message only ever reaches an artifact:
ctest is run with --output-on-failure, so a passing test's output is not in the
job log at all. The suite has been quietly losing most of this test's coverage
in whole configurations and reporting green.

The loop runs in two modes, not one. A window arrives in milliseconds in the
fast mode, and costs between 17 and 40 seconds in the slow one, with nothing in
between across those twenty runs. Ten windows are therefore unreachable inside
any budget worth having: at 40 seconds each that is 400 seconds, and the
unbounded runs measured before any of this took up to 726. Raising the budget
to cover the slow mode would trade a quiet loss of coverage for five
configurations approaching the sixty minute step timeout.

So ask for what a run can reach. Three windows cost 51 to 120 seconds in the
slow mode and under a second in the fast one, and the later hits repeat what
the first ones establish, so what is given up is small. The budget goes to 180
seconds because three windows at the worst rate measured is exactly the 120 it
was, which would have truncated at two.

The count is printed on every run rather than only on a short one. A number
that appears only on shortfall cannot be told apart from a number nobody
recorded.

Reaching the window no times at all is a different matter, and was the worst of
the seven. The check after the loop compares semaphore bookkeeping that a
window has to have touched to mean anything, so a run that reached none of them
compares a counter against the value it was initialised to and reports a pass
having verified nothing. That run now keeps trying to a 300 second ceiling, and
fails if it still has not reached the window. A genuine resonance that holds
for five minutes is worth a failure; the old behaviour was worth nothing.

The SMP copy keeps its count of twenty. It reaches them in under half a second
in all five of its configurations, in all four runs, so the slow mode has never
been observed there and the coverage is free. Both copies get the ceiling and
the unconditional report, so the logic stays identical between them.

Verified locally on all five configurations: the test reaches 3 of 3 in 5 to 14
seconds, and the full suites pass 96 of 96 and 110 of 110 run one test at a
time. With the handler's window made unreachable and the ceiling lowered to 5
seconds, the test stops after 6 seconds, prints the count it reached, and
reports ERROR #8 with the harness recording a failure rather than a pass. With
the count raised past what the budget allows, a run that reaches two windows
still passes, so falling short and reaching nothing stay distinct. The
TX_NOT_INTERRUPTABLE branch, which no configuration in either suite builds, was
compile-checked in both copies with the configurations' own compile commands.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 11:11:31 -04:00
Frédéric DesbiensandClaude Opus 5 c8d27c4e25 Stopped the SMP stack check analyzing a stack it just found broken (#648)
TX_THREAD_STACK_CHECK detects a broken stack, calls the error handler, and then
tests whether the word below the high-water mark still holds the fill pattern.
On the SMP side that second test is a plain if, so a thread whose stack has
just been reported as corrupt goes straight on into _tx_thread_stack_analyze().

Analyzing a stack that is known to be broken is what that function is least
able to do. It binary searches between stack_lowest and stack_highest for the
fill pattern and then scans forward with

    while (*stack_ptr == TX_STACK_FILL)

which has no bound of its own and no reason to terminate once the pattern it is
looking for is no longer where the pointers say it should be. The non-SMP copy
was given an else for exactly this reason. The SMP copy never was, and the two
macros are otherwise identical, line for line, so this single keyword was the
whole of the divergence.

The path is live in CI rather than theoretical. Instrumenting the internal
handler and running all 110 binaries of stack_checking_build shows
threadx_thread_stack_checking_test reaching it four times per run, on a thread
whose stack the test corrupts on purpose. Every one of those four currently
falls through into the analyze it should be skipping.

This is not the timeout the SMP suite has been failing on.
threadx_thread_priority_change never reaches the error handler at all, so
whatever wedges it in teardown is something else. Worth closing regardless: a
runaway scan inside stack analysis would present as a test that stops producing
output and is eventually killed, which is the shape that has been costing this
suite whole runs, and it would be indistinguishable in the log from the hang
already being chased.

Verified on both configurations that define TX_ENABLE_STACK_CHECKING.
threadx_thread_stack_checking_test, the one test that exercises the changed
branch, passes 60 consecutive runs, and stack_checking_build and
stack_checking_rand_fill_build both pass 110 of 110, repeated at the
parallelism CI uses.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 08:52:44 -04:00
Frédéric DesbiensandClaude Opus 5 fe40079353 Stopped the thread priority change test leaving core 0 to a finished thread (#647)
The SMP suite has been failing on threadx_thread_priority_change since 30 June.
The test reports SUCCESS and the process then never exits, so ctest kills it at
the thousand second timeout, twice, and the log carries nothing past the result
line. Instrumenting the harness teardown produced the state at the hang:

  last stage reached:         test_control_return: control thread resume returned
  _tx_thread_preempt_disable: 0
  core 0: current=thread 0 execute=thread 0
  thread test control thread  state=0 priority=0 threshold=0 core_control=1
  thread thread 0             state=1 priority=0 threshold=0 inherit=0

Core 0 is held by a thread in state 1, TX_COMPLETED, while the control thread
sits in state 0, TX_READY, at priority 0. The Linux SMP scheduler fills a core
only when _tx_thread_current_ptr for it is null, and clears that pointer only
for a thread carrying a deferred preemption, which thread 0 is not. So core 0
can never be handed on, and with TX_THREAD_SMP_ONLY_CORE_0_DEFAULT and
TX_SMP_NOT_POSSIBLE the control thread has nowhere else to go. The scheduler
re-reads the same state every two milliseconds for as long as it is allowed to.

What put thread 0 at priority 0 is the last thing this test does:

    thread_0.tx_thread_inherit_priority = 0;
    _tx_thread_smp_simple_priority_change(&thread_0, 16);

with the stated intent of reaching the branch where the new priority is below
the inheritance priority. Zero cannot reach that branch, because 16 is not less
than 0. The other branch runs instead, and that branch assigns the inheritance
priority as the thread's priority while the code after it links the thread into
the list for the new priority regardless. Thread 0 therefore came away claiming
priority 0 while living in the priority 16 list.

Both halves of that hurt. Priority 0 ties with the control thread, so resuming
the control thread raised no preemption and left the execute pointer alone. The
mismatch between the recorded priority and the list the thread is linked into
then means that completing thread 0 removes it from a list it was never in,
leaving core 0 pointing at it for good.

Give the inheritance priority a value above the new one, which is what the
branch the comment names actually requires, and put it back to
TX_MAX_PRIORITIES afterwards so nothing downstream reasons about an
inheritance that is not there. Hold protection across the call as well: this is
an internal routine that expects it, and it was being called in the open.

Measured before and after by printing the thread's state at the point the test
reports success. With the inheritance priority at 0 it is priority 0 threshold
0, matching the hang above, on every run. With it above the new priority it is
priority 16 threshold 16, which agrees with the list the thread is linked into,
and resuming the control thread preempts core 0 the ordinary way.

All five SMP configurations pass 110 of 110 at the parallelism CI uses.

The comparison in _tx_thread_smp_simple_priority_change deserves a second look
on its own account. When its else branch runs, the thread's recorded priority
and the list it is linked into disagree by construction. Only this test is
known to reach that branch, by supplying an inheritance priority that cannot
arise in ordinary operation, so nothing here claims a defect in shipped paths.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 08:52:06 -04:00
Frédéric DesbiensandClaude Opus 5 83dfbc3534 Made a teardown hang in the SMP suite say where it stopped (#646)
The SMP regression suite times out in CI on threadx_thread_priority_change and
the log carries nothing that says why. The reason the log is empty is
mechanical: test_control_return() opens with fflush(stdout), and that is the
last flush before either the test finishes or it wedges. Everything printed
after it sits in stdout's block buffer, and ctest discards that buffer when it
kills the process at the timeout. So the log ends at the test's own result line
no matter where the process actually stopped.

That is enough to place the hang, if not to explain it. The failing runs print
"SUCCESS!" and then stop, which means every check in the test body ran and
passed, and the wedge is somewhere between that flush and exit(). The run of
30 June shows the same signature before any of this test's waits were bounded,
so the hang is not the unbounded wait removed earlier, and the message added
then for an exhausted cap never appears. Retrying tells us nothing new either:
the suite spends 1000 seconds per attempt, twice, to reproduce the same silent
timeout.

Record how far teardown gets, and bound it. A stage variable is updated at each
step from test_control_return() through test_control_cleanup() to exit(), and a
watchdog thread armed on entry to test_control_return() reports the last stage
reached, the per-core scheduler state, and every thread on the created list,
then exits 99.

The watchdog covers teardown and not the test body, because the test body has
no bounded runtime to hold it to. Several tests here wait on a probabilistic
interrupt window: threadx_thread_wait_abort_and_isr_test has been measured
between 0.34 and 439 seconds while passing. Teardown is a fixed amount of work
that takes milliseconds, so a bound on it cannot turn a slow pass into a
failure. The default is 60 seconds, which also means a wedged run now reports
in one minute rather than burning the 2000 seconds two 1000-second attempts
cost today.

The report is written with write() rather than printf() because a wedged thread
may be holding the stdio lock, and a watchdog that blocked on that lock would
reproduce the silent timeout it exists to replace. For the same reason it reads
the ThreadX globals directly and takes no kernel lock; the values may be torn,
which is acceptable for a post-mortem and cannot deadlock.

One walk in test_control_cleanup() is bounded as well. The loop that steps past
the timer thread and the control thread has no terminating condition of its own
and spins for good if _tx_thread_created_count and the created list ever
disagree, which is one of the shapes the timeout could be taking. It now
reports and stops instead.

Off by default in the sense that matters: stderr stays empty and stdout keeps
its buffering, so output is unchanged on a passing run.
TX_TEST_TEARDOWN_TIMEOUT overrides the bound in seconds and zero disables the
watchdog; TX_TEST_TEARDOWN_TRACE echoes each stage as it is reached and
line-buffers stdout so the surrounding output survives a kill too.

Verified against an injected hang at the point the failing runs stop: the
watchdog fires, names the stage, and exits 99. The dump is already informative,
showing thread 0 left at priority 0 with threshold 0 and inherit 0, the same
priority as the control thread, with core 0's execute pointer still on it. All
five SMP configurations pass 110 of 110 at the parallelism CI uses, in both
quiet and trace modes, and the suite runtime is unchanged.

Only the SMP harness is instrumented. The non-SMP suite has not shown this
hang.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 08:51:05 -04:00
Frédéric DesbiensandClaude Opus 5 a5483f0773 Bounded the wait for the delayed suspension window (#645)
threadx_thread_delayed_suspension_test waits for an interrupt to land while
thread 2 is part way through suspending, and waits for it with no bound:

    while(delayed_suspend_set == 0)
    {
        tx_thread_wait_abort(&thread_2);
        tx_thread_relinquish();
    }

How long that takes depends on the build to a degree that is easy to miss. The
loop finishes in between a tenth of a second and three seconds in four of the
five ThreadX configurations. In trace_build it took 490 seconds, which was 40
percent of the whole ThreadX suite and more than every other test in that
configuration put together.

This is the third test in these suites built the same way, after
threadx_thread_priority_change and threadx_thread_wait_abort_and_isr_test: spin
until an interrupt happens to land in a narrow window, with nothing to stop the
spin if it does not. The other two have been given bounds already.

Give this one a wall clock budget too, for the same reason as the last: a tick
is delivered only when the port's timer thread runs, so the tick clock falls
behind real time under load or instrumentation, and instrumentation is exactly
what trace_build turns on.

The check after the loop needs care that the other two did not. It compares
thread_2_counter against thread_2_counter_capture, and the capture is taken
inside the interrupt handler at the moment the window is hit. Leaving that check
in place after a run that never reached the window would compare a live counter
against the zero it was initialised to and report a defect that is not there. So
the check is skipped when the window was not reached, and the run says so.
Reaching the window still exercises it exactly as before.

Verified both ways in trace_build, which is the configuration that was slow: the
window is reached in 8 seconds here and the test passes as it always did, and
with the budget forced to zero the test reports that the window was not reached
and passes without the dependent check firing.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 08:50:15 -04:00
Frédéric DesbiensandClaude Opus 5 4d9ce41845 Gave the wait abort ISR test a budget instead of an open-ended wait (#644)
threadx_thread_wait_abort_and_isr_test waits for an interrupt to land while the
preempt disable flag is set, and waits for it ten times, twenty in the SMP copy,
with no bound on how long that takes. The handler in the same file already says
what can go wrong:

    It is possible for this test to get into a resonance condition in which
    the ISR never occurs while preemption is disabled

and perturbs its own duration to break out of it. That helps but guarantees
nothing, and if the resonance holds, the loop does not end.

It is also, by a wide margin, the most expensive thing in either suite. Run one
test at a time in CI it took between 148 and 726 seconds per configuration:
1936 seconds of the ThreadX suite's 2209, against 273 seconds for the other
ninety five tests together. Nothing else in the suite is within two orders of
magnitude of it.

The budget is in wall clock seconds, not ticks. That distinction turned out to
matter. A tick is delivered only when the port's timer thread gets to run, so
the simulated clock falls behind real time under load or under coverage
instrumentation, and never makes the loss up. A first attempt bounded the wait
at 20000 ticks, nominally 200 seconds, and failed to stop a run that took 726
seconds, because fewer than 20000 ticks had gone by. tx_time_get() cannot bound
elapsed time here; time() can.

A run that falls short says how many windows it reached rather than going quiet,
and the check after the loop is untouched. That check compares semaphore
bookkeeping which holds whatever number of windows were hit, so it still means
exactly what it did before. Hitting the race a few times rather than ten is a
smaller loss than it looks: the value is in reaching the window at all, and the
later hits repeat what the first ones established.

Verified by forcing the budget to 3 seconds, where the test stops after 3.14
seconds of wall clock and reports reaching 0 of 10 windows, with the following
check intact. At 120 seconds both suites pass every configuration run serially.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 08:49:43 -04:00
Frédéric DesbiensandClaude Opus 5 7928dc4268 Ran the ThreadX and SMP regression suites one test at a time (#643)
A large part of both suites sleeps and then asserts on the tick counter, either
exactly or within a tick:

    tx_thread_sleep(18);
    now = tx_time_get();
    if ((now == 18) || (now == 19))

The Linux port drives ticks from a thread running under SCHED_FIFO, so ticks
keep arriving whether or not the thread waiting on them can get a core. Starve
that thread and further ticks land between the sleep expiring and the read, the
value is past the one asked for, and a kernel that behaved correctly is reported
as broken. Thirty-five tests here and forty-one on the SMP side sleep and then
consult the clock or a counter driven by it, so this is most of the suite rather
than a corner of it.

The starvation is self-inflicted. Four tests are run at once on a four vCPU
runner, and each one is a process carrying a scheduler thread, a SCHED_FIFO
timer thread and a thread of its own, so the machine is oversubscribed two or
three times over by design. That is why three of these failed in the run of
18 August, and why the retry hid two of them.

Naming the sensitive tests and keeping them apart was tried first and does not
converge. A list covering the tests comparing tx_time_get() for equality missed
threadx_timer_multiple_accuracy_test, which compares timer-driven counters, and
CI failed on it. Widening the list to cover those missed
threadx_thread_sleep_for_100ticks_test, which asserts a range rather than an
equality, and CI failed on that. A list that is quietly incomplete is worse than
no list, because it reads as protection.

So stop overlapping the tests. Serial execution removes the contention for every
test at once, needs nothing to be enumerated, and makes a run reproducible: a
test either passes on its own machine or has a real defect.

The cost, measured in a four CPU cpuset, is close to a factor of four: the
ThreadX suite goes from 12.2 to 47.6 seconds for a configuration and the SMP
suite from 15.2 to 60.0 seconds. The suites parallelise almost perfectly, so
that factor is what parallelism was buying. It is worth giving up. The run this
replaces spent 36 minutes and reported a timeout carrying no information, and
2000 of those seconds went on retrying a test that had already hung twice.

Serial also makes the tick budget in threadx_thread_wait_abort_and_isr_test mean
what it says. Under contention that test took 255 seconds while its 20000 tick
budget never engaged, because the ticks themselves stretch when the process
cannot get a core. With nothing else running, ticks track wall clock and a
budget in ticks bounds elapsed time.

The FreeRTOS suite is left alone. It covers the creation paths of the
compatibility layer and has no tick accuracy tests, so it has nothing to gain
here.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 08:49:01 -04:00
Frédéric DesbiensandClaude Opus 5 f2d27de25e Stopped a dead apt mirror from taking the whole install down with it (#642)
install.sh reaches the network four times, and on this runner pool that is not
dependable. apt-get update stalled seven times in a single day: once for 55
minutes, once for more than two hours, and five times against the ten minute
step timeout added alongside this. The log says the same thing every time. Every
fetch from azure.archive.ubuntu.com comes back Ign, apt falls back to
archive.ubuntu.com, and then the step produces no further output at all until
something kills it.

Nothing here bounded a fetch and nothing retried one, so a mirror being down
cost a whole run instead of a few seconds.

There is a second problem in the same lines. This script has no set -e, so a
failed apt-get update did not stop the apt-get install that follows. The install
went ahead against whatever package index the image happened to have, and the
run failed later, somewhere with much less to say about why.

Bound each attempt from outside and retry it. apt's own Acquire timeouts were
tried first and are not enough: with them in place a run still sat inside a
single apt-get update for nine and a half minutes without printing a line,
having got as far as fetching noble-security InRelease. The retry loop never got
a turn, because the first attempt never returned, and the step timeout was what
eventually killed it. Whatever apt waits on there is not what
Acquire::http::Timeout covers, so the bound has to come from outside the process.
timeout does not care where the wait is. The Acquire options are kept anyway,
since they make a slow mirror give up sooner, and pip gets its own retry and
timeout flags for the same reason.

timeout goes under sudo rather than over it, so that it signals apt itself.
Signalling sudo risks the kill landing on sudo while apt carries on holding the
dpkg lock, which would leave every retry failing for a different reason than the
one being retried.

The explicit exits stop a failed fetch being carried forward into a build.

set -e is deliberately not used. rm -rf /opt/hostedtoolcache runs without sudo
against a root owned directory and its exit status is not something this script
should start depending on.

The bounds fit inside the ten minute step timeout. Two minutes per attempt,
three attempts, with 10 and 20 second backoffs, caps a command at about six and
a half minutes, and a command that exhausts its attempts exits rather than
letting the next one start.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 08:46:19 -04:00
Frédéric DesbiensandClaude Opus 5 dafb7d70cc Stopped a stalled install from costing a whole regression run (#641)
The Install softwares step runs apt-get update, an apt-get install and a pip
install, and it has stalled twice: 55 minutes on the SMP job of one run, and
more than two hours on the ThreadX job of the next, against a normal 27 to 152
seconds across every other run measured. In the second case the tests never
started at all.

No step in this template had a timeout, so a stall runs until the six hour job
limit. That turns a transient apt or PyPI problem into a lost run, and it hides
what happened: the job simply sits there, and the failure that eventually gets
reported says nothing about which step was stuck.

Bound the three steps that do real work. Ten minutes for the install, against a
normal worst case of 152 seconds. Fifteen for the build, which has run between 4
and 54 seconds. Sixty for the test step, which is the only one whose length
depends on the suites themselves; the longest observed is 37 minutes, and that
was with a wait in one test that has since been bounded.

A step that trips its timeout fails and names itself, which is the point. None
of these numbers is tight enough to trip on work that is merely slow.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 08:44:49 -04:00
Frédéric Desbiens b3486f9b46 Bounded the wait in the thread priority change regression tests (#640)
Both copies of this test, the SMP one and the non-SMP one, install an interrupt
handler and then spin until it clears a flag:

    test_isr_dispatch =  test_isr;
    do
    {
        ...
    } while (test_isr_dispatch);

The handler clears that flag only on a narrow window: thread 3 at priority 6,
ready, and not yet at the head of its priority list, which exists only part way
through a priority change. If an interrupt never lands inside that window, the
loop never ends.

That is what has been failing in CI. The SMP suite has been red since 30 June,
and this test times out in three of the last four failing runs, always in a
stack-checking configuration. The evidence that it is a hang rather than slow
work: the test carries no per-test timeout property, so ctest's --timeout 1000
applies, and locally the test finishes in 0.12 seconds with a worst case of 0.29
over thirty runs. Nothing turns that into more than a thousand seconds. It also
survived --repeat until-pass:2, so it hung twice in succession.

The TX_NOT_INTERRUPTABLE path in the same handler already stops after a fixed
amount of work. Only the interruptable path, which is the one the failing
configuration uses, had no protection.

Cap the loop and clear the handler on the way out. When the window is not reached
the test says so and still passes: not reaching it is a gap in what this run
covered, not a fault in the code under test, and failing would report a defect
that does not exist. The counters are left alone so the checks that follow keep
their previous meaning.

The cap is 100000 attempts. An exhausted cap takes about 60 seconds, measured,
and a successful run takes 0.12 seconds, which puts the usual cost around two
hundred attempts and leaves the cap roughly two orders of magnitude clear of it.
That is wide enough not to lose coverage on a slower machine, while replacing a
timeout that says nothing with a message that says what happened.

Not reproducible here, which fits the diagnosis rather than contradicting it: on
sixteen cores the window is hit almost at once. Thirty sequential runs, two
hundred at parallelism thirty-two, sixty pinned to two CPUs, forty pinned to one,
and three full-suite passes at the parallelism CI uses all came back clean. The
defect is the reliance on the window, not any particular machine.

Verified with the window deliberately made unreachable: before this change the
test runs until it is killed, and after it exits in about a minute reporting that
the window was not reached. The full SMP suite passes 110 of 110 at CI's
parallelism.
2026-08-18 17:01:26 -04:00
Frédéric Desbiens ca62edd27d Swept the stack-heavy measurement across placements, and qualified its result (#638)
#636 reported that a stack in BTCM gave a threefold tighter spread than DRAM0
for stack-heavy work. That measurement used a single code placement, which is
the methodology #631 and #633 exist to correct: the cache benchmark got an
alignment sweep and the interrupt handler got one, and this measurement never
did. It was noticed when #637 added two threads to the same image and the figure
moved -- both spreads came out near 6500 and the minima rose 15%.

The recursive body is now generated at four placements and all four are
measured, per placement, in one image.

    placement    BTCM min / spread     DRAM0 min / spread
    offset 0      41854 / 6850          42036 / 6880
    offset 16     48670 / 1946          48792 / 1978
    offset 32     42388 / 6914          42752 / 7018
    offset 48     48914 / 1860          49198 / 6786

Reproducible across runs to within a few hundred cycles.

Spread is dominated by code placement rather than by the memory holding the
stack. It ranges from 1860 to 6914 depending on where the body falls in a cache
line, and placement also moves the minimum by 17%, from 41854 to 49214. Against
that, the memory contributes a consistent but small advantage: BTCM's minimum is
lower at all four placements, by 0.4% to 0.9%.

BTCM's spread beats DRAM0's decisively at one placement of the four, offset 48,
at 1860 against 6786. At the other three the two are within 2% of each other.
So the effect #636 reported is real where it occurs and is not a property of the
part: quoting it as one invited the reader to expect it everywhere.

#636's claim should be read as qualified by this. A stack in BTCM buys a small
consistent improvement in the best case and a large improvement in spread at
some code placements and not others. Anyone building a determinism argument on
it needs the placement sweep in the loop, not a single figure.

The pad nops that displace each placement execute on every recursion level
rather than once, so each placement carries a slightly different constant cost,
about 0.6% at the widest. That cancels in the BTCM against DRAM0 comparison,
which is made at the same placement, and does not affect spread within one.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-18 08:51:49 -04:00
Frédéric Desbiens 1ef49338e9 Gave each thread its own MPU window, and made a violation fault (#637)
First step towards a ThreadX module port for this core: establish that PMSAv8-R
regions can be switched per thread on this part, what that costs, and that a
violation actually faults. Those are the questions worth answering before
writing a module manager on top of them.

Two threads each own a 4 KB window at the top of DRAM2. The windows are carved
out of the broad data region in mpu.c, because isolation is only meaningful in
memory no other region already covers -- every other region in that map is a
wide RW window, so a private buffer inside one of them would be reachable by
every thread whatever else was programmed.

Each thread writes its own window, which must succeed, and then reaches for the
other thread's, which must fault. The second half is the part that matters: a
test that only shows a thread reaching its own memory would pass just as well
with no protection at all.

Measured on the S32Z280-594EVB, reproducible across three runs:

    thread 0 window 0x3187E000  own: reachable  other: faulted
    thread 1 window 0x3187F000  own: reachable  other: faulted
    region switch cost: 562 to 604 cycles

The cost is worth noting for the module port to come. A context switch on this
part is about 1400 cycles, so switching one region adds roughly 40% to it, and
most of that is the dsb and isb rather than the register writes. A module switch
programming several regions should therefore batch the barriers once at the end
rather than per region.

Scope, stated plainly. The window is applied by the thread calling
thread_mpu_activate, not by the scheduler. The port's scheduler does call
_tx_execution_thread_enter under TX_ENABLE_EXECUTION_CHANGE_NOTIFY, which would
make it automatic, but that macro is read by port assembly compiled into the
shared threadx library, so enabling it would oblige all nine example targets in
this port to supply the four execution hooks. A ThreadX module port carries its
own copies of the port assembly for exactly that reason, and that is where the
switch belongs. There is no user mode, no syscall boundary and no loader here.

The fault is survivable the same way the boot probes make it survivable:
fault_expected tells the data abort handler to record the violation and resume
after the faulting access. That works in thread context because the handler
returns where it came from rather than to a fixed recovery point.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-17 17:59:12 -04:00
Frédéric Desbiens f535a4ec67 Measured stack-heavy work against the memory holding the stack (#636)
#635 found BTCM worth about 7.4% on a context switch with no determinism
advantage, and said why: a switch saves sixteen registers, roughly one cache
line, so the stack's cache state has almost nothing to contribute. It named the
interesting case as work with a large stack working set, said it had not been
measured, and said it should not be assumed. This measures it, and the answer
inverts the earlier one.

deep_touch recurses 24 frames, writing a frame on the way down and reading it on
the way up, so the working set is the whole descent. The cache is cleaned and
invalidated before each sample, so every descent starts cold. Two threads, one
stack in BTCM and one in DRAM0, no partner threads and no relinquish: the timed
region is entirely within one thread.

Reproducible across runs:

                      min      mean      max     spread
    stack in BTCM   42868     43031    44890       2020
    stack in DRAM0  42982     43291    49842       6860

The mean is the same to within 0.6%. The worst case is 10% lower for BTCM and
the spread is 3.4 times tighter. No sample in either configuration exceeded
twice the minimum, so these maxima are the workload rather than a timer tick --
which is the mistake that produced a false jitter result in #635 and is why the
count of interrupted samples is printed.

Put beside #635 the two measurements say opposite things and both are true.
For a context switch, a small footprint touched every time, BTCM buys throughput
and no determinism. For stack-heavy work, a large footprint touched once, it
buys determinism and almost no throughput.

The reason is that this workload is compute bound at the optimisation level this
BSP builds at: 43000 cycles for 24 frames is dominated by call and loop overhead,
so line fills are a few percent of the total and barely move the mean. What they
do is vary, and that variance is what a bank with no cache in the path removes.

So the determinism argument for TCM holds here, but it is worth 10% of worst case
and a threefold narrowing of spread, not an order of magnitude. Anyone citing
this in a safety argument should cite those numbers and not a larger claim.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-17 17:26:08 -04:00
Frédéric Desbiens 367d91880b Measured context-switch cost against the memory holding the stack (#635)
#634 placed a thread stack in BTCM and deliberately claimed no timing benefit,
because none had been measured. This measures it.

Two pairs of equal-priority threads hand control back and forth with
tx_thread_relinquish. One pair has both stacks in BTCM, the other in DRAM0, and
the measuring thread of each pair times the round trip in PMU cycles. Both pairs
run in one image from one copy of the measuring code, which is what makes the
comparison safe: the alignment trap that invalidated earlier work here bites
when two builds with different layouts are compared, and a code shift moves both
pairs equally.

Reproducible to the cycle across runs:

                        min    mean    max
    BTCM stacks        1370    1379   1402
    DRAM0 stacks       1476    1480   1508

BTCM is about 7.4% faster, or 110 cycles on a round trip of two switches.

Three findings that bound the claim, and the last one deflates it.

The figure holds whether the cache is warm or cold. Cleaning and invalidating
the data cache before every timed switch costs both configurations about 40
cycles and leaves the gap at 7.4%: warm it is 1334 against 1440, cold 1370
against 1476. So the advantage comes from BTCM's zero wait states, not from
avoiding cache misses.

That is because a context switch touches almost no stack -- sixteen registers,
about one cache line -- so the stack's cache state has little to contribute
either way. TCM should matter much more for threads with deep call chains or
large locals, where the stack working set is big enough for cache state to
dominate. That is not measured here and should not be assumed.

There is no determinism benefit visible in this test. Excluding preempted
samples, jitter is 32 cycles for BTCM and 28 to 34 for DRAM0 -- comparable, not
better. A first version of this measurement appeared to show BTCM with 16 times
less jitter, and that was wrong: max was reporting whichever pair a timer tick
had landed on. Across three runs the outlier appeared in the BTCM pair once and
the DRAM0 pair twice. Samples past 2000 cycles are now counted separately and
excluded from min, mean and max alike, and the count is printed so the reader
can see how many there were.

Also tried and discarded: loading the partner thread with a cache walk to create
pressure. The timed round trip includes the partner, so the walk dominated every
sample and put all 256 past the outlier threshold. The per-sample flush replaced
it and sits outside the timestamps.

The demo also starts the PMU cycle counter, which bsp_boot.c does for the probe
image and this image never ran. Without it every reading would have been zero,
which reads as a free context switch rather than as a dead counter.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-17 17:08:29 -04:00
Frédéric Desbiens 62e966f6bf Enabled BTCM, measured it, and put a ThreadX thread stack in it (#634)
* Enabled BTCM and measured what it offers as a data store

BTCM was disabled because of a regression that turned out not to exist; the
claim was retracted in the previous commit. It is enabled now, and this
measures why that is worth doing: 16 KB at zero wait states, where ATCM has
one, and no cache in the path at all.

Enabling it needs three things, all of which existed for ATCM already: the
region register write at EL2, an MPU region, and an ECC preload before any read
(TRM 6.2.2). BTCM accepts 32-bit stores where ATCM needs 64-bit, which
tcm_preload already handles.

The cache benchmark now sweeps three memories rather than one, four loop
alignments each. At the alignments where the loop is not instruction-fetch
bound:

    memory              cold (uncached)   warm (cached)   gain
    DRAM2 half-speed            857,540         651,436   24.0%
    DRAM0 full-speed            797,824         651,297   18.3%
    BTCM  zero wait             694,689         651,369    6.2%

Three things follow.

Warm times are identical across all three memories, within 0.02%. Once the data
cache is working the backing store barely matters, because the working set fits
in it.

Cold times rank as the reference manual predicts: BTCM fastest, then DRAM0,
then DRAM2 at half the core frequency (S32Z2 RM 6.3.6).

BTCM still shows a 6.2% gain when the caches are enabled, and that cannot be
the data cache, because an enabled TCM is Non-cacheable Non-shareable Normal
memory whatever the MPU says. It is the instruction cache on the timing loop.
This probe has always measured both caches together; three memories side by
side is what makes that visible.

The number that matters for placing data in BTCM: uncached BTCM is within 6.6%
of the best cached case, where uncached DRAM0 is 22% off it. Data in BTCM runs
at close to cache-hit speed with no cache to miss, which is the determinism
argument stated as a measurement rather than an assertion.

DRAM2's sweep is unchanged with BTCM enabled, 0 and 0 and 240 and 240, which
independently confirms the retraction.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Put a ThreadX thread stack in BTCM

The code side of TCM was done in #630; this is the data side. One of the demo's
three thread stacks now lives in BTCM and the other two stay in DRAM0, so a run
exercises both paths and a mistake in either shows up.

link.lds gains a BTCM region and a .btcm_bss NOLOAD section, so a stack is an
ordinary C array with a section attribute and the linker checks it fits, rather
than a hardcoded address that silently overflows the bank.

entry.S preloads the whole bank at EL2, and that is not optional. ECC is enabled
on this part, so a TCM location must be written before it can be read (TRM
6.2.2), and a stack is read before the program writes it -- the first context
restore pops what tx_thread_create built into it. The preload has to happen
before any C runs, because the demo images do not run bsp_boot.c, which is where
the ATCM preload lives. 32-bit stores suffice for BTCM where ATCM needs 64-bit.

Why BTCM for a stack: 16 KB at zero wait states where ATCM has one, and never
cached whatever the MPU says about it. Measured in the previous commit, uncached
BTCM comes within 6.6% of the best cached case while uncached DRAM0 is 22% off
it, so stack access runs at close to cache-hit speed without depending on a line
being resident. That is the property a determinism argument needs.

What this commit does not claim: no thread-level timing improvement has been
measured. The case for BTCM here rests on the memory characterisation and on
removing the cache from the path, not on a measured context-switch figure. That
measurement is worth doing and has not been done.

Verified on the S32Z280-594EVB. The demo reports its stack addresses so the
placement is visible rather than implied -- sleeper at 0x30100000 in BTCM,
spinner and judge in DRAM0 -- and passes with 100 ticks, 20 sleeper wakeups and
20 preemptions, so a real thread schedules, preempts and context-switches on a
tightly-coupled-memory stack. The boot image still passes six of six probes.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-17 16:33:55 -04:00
Frédéric Desbiens 35de56853c Retracted the claim that a second TCM bank costs the data cache (#633)
entry.S and readme_s32z280.txt both stated that enabling any second TCM bank
removes all measurable data-cache benefit on this part, and gave five
configurations as evidence: ATCM alone at a 24% cache gain, four combinations
involving a second bank at none. Both concluded the cause was not a particular
bank, not its address and not the ECC preload, but enabling a second bank at
all. Both cited the Cortex-R52 and S32Z2 errata as not covering it and offered
a shared RAM pool between the LLC and the TCMs as an explanation. A defect
report went to NXP on that basis.

It was an artifact of the benchmark. That benchmark was bimodal with respect to
where its timing loop fell inside a 64-byte cache line, reporting either 24% or
nothing at all for identical silicon, and every one of those five
configurations was an edit to entry.S, so every one shifted the code that
followed and moved the loop between modes. Adding two nop instructions
reproduces the "second bank" figure exactly, to the digit. The report to NXP
has been withdrawn.

Re-measured with the alignment sweep added in #631, one bank and two are
indistinguishable:

    loop offset in line      ATCM only      ATCM + CTCM
    0                        gain 0         gain 0
    16                       gain 0         gain 0
    32                       gain 240/1000  gain 240/1000
    48                       gain 240/1000  gain 240/1000

So enabling a second bank costs nothing measurable. The banks stay disabled,
but for the ordinary reason that nothing in this example uses them, and both
texts now say that instead. Enabling one is a single line, with the ECC preload
before any read (TRM 6.2.2) and an MPU region as the only prerequisites, both
already handled for ATCM.

The readme also now states the general point, which outlasts the TCM detail: a
single-figure timing result from this example cannot be compared across builds
unless the timed loop's alignment is controlled, because almost any change
shifts code.

Comments and documentation only; no generated code changes. Verified on the
board regardless, since entry.S was touched: six of six probes pass and the
sweep is unchanged.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-17 13:46:33 -04:00
Frédéric Desbiens 3aaaa6700d Revalidated the ATCM handler result across four placements, and it holds (#632)
The handler comparison in #630 measured one alignment, which is the mistake
that made the cache benchmark in this example report 24% or 0% for identical
silicon. The handler body is now generated at four offsets within a cache
line, all four are measured, and the figures are reported per placement.

The claim survives. Mean cycles for the handler body:

    loop offset      code RAM    ATCM
    0                     523     383
    16                    534     377
    32                    539     375
    48                    521     383

ATCM is faster at every placement, by about 28%, and the two sets of means do
not overlap. Worst case improves as well, 510 against 694.

Two things worth recording beyond the headline.

The handler measurement is only mildly alignment sensitive, 3.5% across
placements in code RAM and 2% in ATCM, quite unlike the cache loop's two
modes. So this comparison was less fragile than the cache one, and #630's
direction was right even though its method was not defensible. The absolute
numbers differ from #630 because the body now sits behind a placement wrapper
that adds a call; the comparison is internally consistent either way.

Both variants also report identical cache sweeps, 0 and 0 and 240 and 240,
which settles the regression this branch's predecessor appeared to show. That
apparent regression was the single-alignment probe moving between its two
modes, not anything about ATCM.

One copy of the logic is kept: the wrappers inline a single always_inline
implementation, so the four placements cannot drift apart.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-17 13:33:43 -04:00
Frédéric Desbiens 1d3f3f8e4c Measured the cache benchmark at four alignments, because one is not enough (#631)
This benchmark was bimodal and reported a single number, which made it worse
than no benchmark. The same workload on the same silicon reports either 24%
cache benefit or none at all, decided only by where the loop falls inside a
64-byte line -- and therefore by any unrelated change that shifts code
ahead of it. Two nop instructions added to entry.S were enough to flip it.

That is not a hypothetical. A run of conclusions drawn from this probe turned
out to be measuring code layout: an interrupt-handler comparison, a claim that
enabling a second TCM bank costs all cache benefit, and a follow-up claim that
what mattered was when the TCM region register was written rather than what it
contained. The last of those was reported to NXP as a defect and has had to be
withdrawn. Enabling CTCM and adding two nops produce identical results, to the
digit, because the only measurable consequence of the enable was the eight
bytes of instructions it added.

Four copies of the loop are now generated at different offsets within a cache
line, all four are measured, and the low and high gains are both reported.
Pinning a single alignment was tried first and is not a fix: it silently picks
one of the two modes -- aligned to 64 the loop sits permanently in the low one.

Measured on the S32Z280-594EVB, reproducing exactly across runs:

    loop offset in line    cold      warm      gain
    0                      890,302   890,035   0%
    16                     890,208   889,976   0%
    32                     857,439   651,390   24.0%
    48                     857,631   651,571   24.0%

The cold pass differs between the modes as well, 890k against 857k, so the
loop is slower even with both caches off. The cold pass is instruction-fetch
bound out of code RAM at half the core frequency (S32Z2 RM 6.3.6), and how the
loop straddles lines decides how much of the data cache's contribution is
visible at all. This probe therefore measures both caches together and always
did; the sweep at least makes the variation visible instead of letting one
arbitrary placement stand in for the part.

C4 now passes if any alignment shows a 10% speedup, and says so explicitly
when the low mode does not, so the sensitivity appears in the log rather than
being discovered later.

Verified: with the sweep in place, adding 0, 8, 12 or 20 bytes of nops to
entry.S leaves the reported low and high gains unchanged. Before it, the same
shifts read 24.0%, 0%, 0% and 0%.

Also adds cache_disable_all, which the sweep needs: cache_enable was one-way,
so a second cold reading in one run was impossible.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-17 13:14:16 -04:00
Frédéric Desbiens b2983c08d3 Ran the interrupt handler from ATCM, and measured what that buys (#630)
The TCM work so far enabled ATCM and left it empty, which buys nothing.
This places code in it and measures the result.

link.lds gains an ATCM region and an .atcm_text section whose run address
is in the bank and whose load address is in CODE. tcm_copy_atcm_text
moves it, using 64-bit stores because ECC is enabled on this part and
ATCM requires them (Cortex-R52 TRM 6.2.2); both ends of the section are
8-byte aligned so there is no narrower tail to leave without check bits.
The copy runs after T4 and T5, which write test patterns to the first and
last words of the bank and would otherwise land on top of the code.

s32z280_atcm.elf is the same image as s32z280_boot.elf with the interrupt
service body placed in ATCM. Both targets exist so the comparison can be
repeated on one board in one session without reconfiguring. The service
routine is split into a timed wrapper that stays in .text and a body that
moves, so the wrapper's own cost appears in both measurements and cancels.

Measured in PMU cycles over 64 samples, caches enabled in both:

                 code RAM    ATCM    change
    min               454     334    -26.4%
    mean              458     340    -25.8%
    max               612     466    -23.9%
    spread            158     132    -16.5%

CNTPCT is not used for this: at 8 MHz it cannot resolve a handler body,
let alone the variation in one.

The level shift is the solid part. ATCM is a quarter faster even though
the caches were on and code RAM had the instruction cache available,
which says the handler does not stay resident between interrupts 10 ms
apart -- so each one pays a cold fetch from code RAM, which runs at half
the core frequency where ATCM runs at full speed with one wait state
(S32Z2 RM 6.3.6).

The determinism claim deserves less weight than the numbers first
suggest. The spread narrows by only 16%, and ATCM's worst case still sits
slightly above code RAM's best case, so the two distributions overlap at
the tails rather than separating. Whatever jitter remains is not
dominated by instruction fetch.

Both images pass six of six boot probes.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-17 09:21:27 -04:00
Frédéric Desbiens ff638dae71 Configured the LINFlexD console twice, because once is not enough at -O2 (#629)
The console driver runs its configuration sequence a single time, and that
works only because this BSP is built without an optimisation flag. Compiled
at -O2 the same sequence leaves the line corrupted: every character
partially wrong, in the pattern the file already describes for a
misconfigured module.

What makes it worth guarding against is that the failure is invisible.
UARTCR reads back exactly the value written. LINIBRR and LINFBRR read back
exactly the values written. linflexd_init returns LINFLEXD_INIT_OK. The
registers are right and the line is wrong, so nothing in the returned
status tells the caller the console cannot be trusted.

Localised by bisection: with every other file at -O2 and this one at -O0
the output is clean, and with only linflexd_init at -O2 it is corrupted,
so the fault is in the configuration sequence rather than in the per-byte
transmit path.

The mechanism is not understood, and this commit does not claim to explain
it. Tested and rejected: a 100x larger bound on the wait for
initialisation mode, a settling delay before the first LINSR read, a
settling delay after leaving initialisation mode, a barrier and read-back
between the two UARTCR writes, and waiting for LINSR to report the exit
from initialisation mode. None of those makes a single pass work at -O2.
A second pass does, at both optimisation levels, which is what this does.

Instrumented with a duplicate of the sequence forced to -O2 and reported
through a console repaired afterwards, which is how the register read-backs
above were obtained.

Verified on the S32Z280-594EVB. The boot image passes six of six probes
with the console status still reporting 0x00000000, and the reproducer
builds clean and prints correctly at both -O0 and -O2.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-17 07:59:23 -04:00
Frédéric Desbiens 2f6945475b Stopped pthread_self() faulting when the caller is not a pthread (#627)
Nothing prevents an application mixing tx_thread_create() with the POSIX layer,
and a thread created that way has no POSIX control block. posix_thread2tcb()
returns NULL for it, which posix_thread2tid() then read through:

    p_tcb = posix_thread2tcb(thread_ptr);
    thread_ID = p_tcb->pthreadID;

pthread_self() went on to compound it, reading the signal fields of a POSIX_TCB
out of a thread that is only a TX_THREAD:

    if (((POSIX_TCB *) thread_ptr) -> signals.signal_handler)

The first is a null dereference and the second runs off the end of the control
block into whatever the linker put there. Under qemu-system-riscv32 the first one
lands first: mcause=0x5, a load access fault, with mtval=0xb4 for the offset of
pthreadID.

Have posix_thread2tid() report zero for a thread with no control block, which is
what px_pth_join.c already does for the same call, and have pthread_self() skip
the signal check unless the ID says the caller really is a pthread. Zero cannot
collide with a real ID because px_pth_create.c uses the address of the control
block as the ID.

This also covers the case where there is no current thread at all, from an ISR or
before the scheduler starts: tx_thread_identify() returns NULL, and the same
zero comes back instead of a fault.

Add posix_pthread_self_test, which asks both kinds of thread for their ID: a
pthread, which has to report what pthread_create() returned, and a plain ThreadX
thread, which has to report zero. Reverting either half of the fix turns the test
into the load access fault above.

Verified with riscv64-unknown-elf and qemu-system-riscv32: 4 tests across the
default build, 4 of 4 passing.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-16 18:45:16 -04:00
Frédéric Desbiens f184531e7f Added the POSIX compatibility layer to the CMake build, with regression tests (#626)
* Added the POSIX compatibility layer to the CMake build

Nothing in the repository built the POSIX layer. The FreeRTOS layer next to it
has had a target since the CMake build was introduced, so the POSIX one was the
odd one out, and 106 source files went unbuilt by any target, on any
architecture.

Add a posix-threadx target, following the FreeRTOS layer's shape: a static
library, EXCLUDE_FROM_ALL so the default build is unchanged, linking threadx and
publishing its own directory as a PUBLIC include path. The sources live in their
own CMakeLists.txt rather than the top-level file, as common/ does, because there
are 106 of them. The seven posix_*.c files in the same directory are a demo and
standalone signal tests, each with its own entry point, so they stay out of the
library.

The layer does not suit every configuration, and the target is only offered where
it can work:

  - Hosted simulation ports (linux, win32, win64) build against a C library that
    already provides errno.h, pthread.h and the rest. The layer replaces those.
    Its pthread.h even uses _PTHREAD_H, the same include guard as glibc's, so its
    declarations are skipped wholesale and the build fails on missing types.
    Neutralising that guard only exposes the real problem: 69 conflicting
    definitions in a single translation unit, for time_t, struct timespec,
    sigset_t, pthread_t, pthread_mutex_t, sem_t and more. Both the layer and the
    C library implement POSIX, and only one of them can define those names. The
    linux port also emulates threads by calling the C library's pthread_create
    and sem_wait, which the layer exports itself, so linking the two would divert
    the port into the layer that sits on top of it.
  - SMP builds. px_int.h declares _tx_thread_current_ptr as a plain pointer,
    which is a per-core array under SMP, and the layer tracks no current core.

Building the layer for the first time exposed one portability defect worth
fixing rather than working around. tx_posix.h defined ssize_t as INT, with a
comment conceding it should come from <sys/types.h>. That is correct only where
the C library agrees: on AArch64 newlib makes ssize_t 64 bits, and every
translation unit that reached a library header failed to compile. Defer to the
library when it has declared the type, keyed on the _*_DECLARED guards newlib
uses, and do the same for mode_t, which had the same problem waiting. Where no
library declaration exists the previous definitions still apply, so the 32-bit
targets that did build are unaffected.

Verified by building posix-threadx for arm9, arm11, cortex_m0, cortex_m3,
cortex_m4, cortex_m7, cortex_m33, cortex_m55, cortex_m85, cortex_a7, cortex_a9,
cortex_r4, cortex_r5, cortex_a34, cortex_a53 and cortex_a55 with
arm-gnu-toolchain-14.3.rel1, and for risc-v32 and risc-v64 with
riscv64-unknown-elf: 18 of 18, 106 objects each. cortex_a78 has no non-SMP port
and fails to configure with or without this change. Linking the result against
libthreadx.a leaves only tx_application_define, _tx_initialize_low_level, the
optional execution profile hooks, and memset and strlen unresolved, all of which
the application or its C library supplies. The default build still produces
libthreadx.a and no POSIX library.

Compiling is not the same as working, and on the 64-bit targets in that list it
is not enough. The layer carries a message by putting the address of a private
buffer into the queue, and ULONG is 32 bits on every port, so that address only
fits when TX_64_BIT is defined. Without it px_mq_send.c truncates the pointer
and px_mq_receive.c casts the truncated value back, which GCC reports as nothing
worse than a -Wpointer-to-int-cast warning. Defining TX_64_BIT is not a remedy
either: tx_api.h then reaches for the extension pointer macros, which need
tx_thread_extension_ptr in the thread control block, and outside ports_smp and
ports/linux no port declares it. So the target builds everywhere, but the
message queues are only sound on the 32-bit ports. That is pre-existing, it is
not made worse here, and it is left for a change of its own.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Added regression tests for the POSIX compatibility layer

The POSIX layer had no tests. The seven posix_*.c programs shipped beside it are
demos: they print nothing, report no result and end in infinite loops, so they
tell a person watching a debugger something and an automated run nothing.

Add a suite under test/posix, laid out like the FreeRTOS one and driven the same
way, with scripts/build_posix.sh and scripts/test_posix.sh over a run.sh that
takes the same arguments as its RISC-V counterpart.

The tests run on emulated hardware because they have nowhere else to go. The
layer replaces the C library's POSIX headers and exports the same symbols the
linux port calls to emulate threads, so a host build is not available to it. The
RISC-V QEMU harness that the ThreadX suite already uses is, and this suite reuses
its BSP and testcontrol.c rather than growing copies of them.

Three tests to start:

  - posix_mq_basic_test sends a message through a queue and checks the contents
    and priority survive the round trip.
  - posix_mq_send_abort_test covers the leak fixed in #624, by filling a queue,
    blocking a sender on it, aborting the wait and watching the queue's byte
    pool. Reverting the fix makes it fail on the pool check, so it measures what
    it claims to.
  - posix_pthread_basic_test covers pthread creation, a mutex, a semaphore
    handoff, pthread_self and collecting an exit value through pthread_join.

The queue's pool is sized (mq_maxmsg + 1) * (mq_msgsize + 11), which leaves room
for about one message beyond a full queue, so the abort test uses small messages
and a shallow queue. With a larger message the first leaked buffer exhausts the
pool, tx_byte_allocate fails, and the sender disappears into the endless loop in
posix_internal_error() instead of reporting anything. Sizing it this way keeps
the failure legible as a pool measurement rather than a timeout.

riscv32 only, and the reason is the layer rather than the harness. The layer puts
the address of a message buffer into the queue, ULONG is 32 bits on every port,
and a 64-bit address only fits there when TX_64_BIT is defined. Defining it makes
tx_api.h use the extension pointer macros, which need tx_thread_extension_ptr in
the thread control block, and no port outside ports_smp and ports/linux declares
it. Configuring for risc-v64 stops with that explanation rather than building
something that would corrupt a pointer at runtime.

Verified with riscv64-unknown-elf and qemu-system-riscv32: 3 tests across
default_build, disable_notify_callbacks_build, stack_checking_build and
trace_build, 12 of 12 passing.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-16 18:24:48 -04:00
Frédéric Desbiens 3501c5c31d Stopped two POSIX call sites falling through their error handler (#625)
posix_internal_error() never comes back for a non-zero code: its body is
"while (error_code) { ; }". Most callers in the layer still follow it with an
explicit error return. Two do not, and both fall through into a pointer
dereference, so they are only correct because the handler happens to hang.

That is a fragile thing to rely on. The C standard permits an implementation to
assume a loop with no side effects terminates (C11 6.8.5p6), so the guarantee is
a property of the toolchain rather than of the language. Both toolchains the
project supports do preserve the loop today - checked with GCC 14.3 at -O0, -O1,
-O2 and -Os, and with Clang 22 at -O0 and -O2 - so nothing is broken right now.
Neither call site should depend on that.

mq_send() falls through with bp indeterminate, having just been told the
allocation failed, and would copy msg_len bytes through it. Report ENOMEM and
return ERROR instead.

posix_thread2tid() falls through with thread_ptr NULL. posix_thread2tcb() returns
NULL for that input, and the next line reads p_tcb->pthreadID. Return zero, which
is never a valid pthread ID because px_pth_create.c assigns the address of the
TCB as the ID.

No behaviour changes while the handler keeps hanging; both additions are
unreachable today.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-16 18:24:25 -04:00
Frédéric Desbiens 2f378a4c8e Stopped POSIX mq_send() leaking the message buffer when the send fails (#624)
mq_send() allocates a private buffer from the queue's own byte pool, copies the
caller's message into it, and passes the buffer's address through the queue. The
receiver takes ownership: px_mq_receive.c releases the buffer once it has copied
the message out.

If tx_queue_send() fails, the message never reaches the queue, so no receiver
will ever release that buffer. mq_send() returned ERROR while still holding the
only pointer to it, leaking it from the pool.

The failure is reachable. With TX_WAIT_FOREVER the send suspends, and
tx_queue_send() then returns the thread's suspend status. tx_thread_wait_abort()
sets TX_WAIT_ABORTED on a suspended sender, and the queue survives that, so the
pool keeps shrinking with every aborted send.

Queue deletion also reaches the branch, via TX_DELETED, but vq_message_area is a
TX_BYTE_POOL embedded in the queue structure and destroyed with it, so nothing
outlives the failure there.

Exhausting the pool does not merely make later calls fail. tx_byte_allocate()
failure runs into posix_internal_error(9999), which busy-loops forever on a
non-zero code, so a caller hangs rather than getting an error back.

Release the buffer before returning. The EINTR reporting is unchanged, since
TX_WAIT_ABORTED maps onto it reasonably.

Reported-by: K-ANOY <https://github.com/eclipse-threadx/threadx/issues/568>
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-16 09:45:23 -04:00
Frédéric Desbiens 0e7db2bf79 Stopped the generated Armv8-M readmes claiming a false origin date (#623)
Every Armv8-M port readme ends with

    09-30-2020  Initial ThreadX 6.1 version for Cortex-M85 using GNU tools.

with the core name substituted in by scripts/copy_armv8_m.sh. The date is the
shared port's, so each core inherits it whatever its own history: Cortex-M85 was
announced in 2022 and its readme claims a 2020 origin, and any core added later
gets the same treatment the moment its name joins the generator's list.

The rest of the history block is accurate, since it records changes to the shared
files. Only the closing line asserts something per-core. Reword it to describe
the Armv8-M port itself, and say where a given core's real starting point is.

Regenerating updates the twelve readmes for cortex_m33, cortex_m52, cortex_m55
and cortex_m85 across the three toolchains.

Cortex-M52 makes the point: it arrived in #519 and its readme immediately claimed
a 2020 origin for a core announced in 2023.

The ARMv7-M templates say "Initial ThreadX version 6.1.7 for Cortex-M", with no
placeholder to substitute, so they make no per-core claim and are left alone.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-16 08:42:53 -04:00
Frédéric Desbiens 09c190a71d Added a CMake target for the Linux sample program (#622)
The CMake build produced libthreadx.a and nothing else, so trying ThreadX on
Linux meant using the Makefile beside the port instead. Build the demo the
Makefile builds, for the linux port and its SMP counterpart.

The target is behind an option that defaults off, so an ordinary build is
unchanged and still produces just the library. -DTHREADX_SAMPLE=ON adds it:

    cmake -S . -B build -DTHREADX_ARCH=linux -DTHREADX_TOOLCHAIN=gnu \
          -DTHREADX_SAMPLE=ON
    cmake --build build --target sample_threadx

The include path uses TX_COMMON_DIR rather than naming common or common_smp,
since the top level already resolves which of the two applies.

Verified by building and running both variants. Non-SMP prints

    **** ThreadX Linux Demonstration **** (c) 1996-2020 Microsoft Corporation

and SMP prints the SMP banner, both with the demo's thread counters advancing. A
default configure with no THREADX_SAMPLE has no sample_threadx target and still
produces libthreadx.a, so nothing existing moves.

Derived from the two example_build files in #404 by Yanfeng Liu, which had the
same goal. That change also rewrote the top level's SMP selection, added
common_smp/CMakeLists.txt and added ports_smp/linux/gnu/CMakeLists.txt; all three
have since arrived on dev by other routes, so only the sample targets were still
missing. The include path needed adjusting because the original depended on
THREADX_SMP being a string suffix, which it no longer is.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-15 17:28:33 -04:00
Frédéric Desbiens 1296cf1740 Corrected the S32Z280 data SRAM map against the Reference Manual and the part (#621)
The RTU has 1 MB of data SRAM in three contiguous banks, and this port
described it wrongly in both directions.

  DRAM0  0x31780000  256 KB  full core speed
  DRAM1  0x317C0000  256 KB  full core speed
  DRAM2  0x31800000  512 KB  half core speed

DRAM1 was not declared at all, so 256 KB of full-speed memory went unused
and the linker's DATA region stopped at 256 KB. DRAM2 was declared as
2 MB when it is 512 KB, which mattered more: the MPU mapped 1.5 MB past
the end of the bank, and that range aliases back onto its base. Anything
placed above 0x31880000 would have shared storage with the bottom of the
region silently -- no fault, two objects at one address. Nothing was
placed there yet, so this was a trap rather than a live defect.

Sources: S32Z2 Reference Manual Rev. 5, section 6.3.6 and Table 13, and
the board. Writing distinct values to all three banks and reading them
back shows 1 MB of independent storage, and the first word past DRAM2
returns the value written to its base, which is what fixes the size.

Note that NXP's own debugger memory map, s32z2e2_memory_regions.py in
S32 Design Studio, calls the last bank 2 MB. The Reference Manual and the
silicon agree it is 512 KB.

Also corrected the description of these banks throughout. They are all
RTU-local; the earlier comments treated locality as the thing that
distinguishes them, when the actual difference is clock speed. That is
why the cache benchmark uses DRAM2 -- caching a bank that already runs at
core speed shows nothing, which is a real effect the old wording
explained with the wrong cause.

Verified on the S32Z280-594EVB: builds clean, and the boot probes pass
six of six with the MPU, GIC, interrupts, caches and both protection
faults exercised.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-15 17:02:01 -04:00
Frédéric Desbiens 358e7a9ea5 Stopped the remaining example scripts naming archives that were deleted (#620)
#618 fixed the arm9 and arm11 shell scripts, but not their .bat counterparts,
and did not look at cortex_r4 and cortex_r5 at all. An audit of every build
script against the files actually present in its directory found the rest.

The .bat scripts for arm9, arm11, cortex_r4 and cortex_r5 still linked libc.a,
libgcc.a and, for arm11, libnosys.a from their own example_build directories,
and the cortex_r4 and cortex_r5 shell scripts did too. Those archives went in
6.1.10 under "Removal of unneeded files", so on Windows all four examples failed
exactly as the shell versions did before #618, and on Linux the two R-profile
ones still did.

Link through the compiler driver, as the other examples have since #594.

This does not make the examples link, and the change stops there deliberately.
All four now fail the same way,

    undefined reference to `_fini'

because their linker scripts define the .init and .fini sections but not the
_init and _fini symbols, which live in crti.o and crtn.o and are omitted by
-nostartfiles. Reviving four very old cores is separate work.

Correct the comment on EXAMPLES_EXPECTED_TO_FAIL again. It had cortex_r4 and
cortex_r5 failing for want of newlib multilib variants; they do not. All four
share the single cause above, and the multilib explanation was wrong for the
R-profile pair just as it was for arm9 and arm11.

Verified by running each shell script: cortex_r4 and cortex_r5 fail on _fini
rather than on missing files, matching arm9 and arm11. The .bat changes mirror
link lines proven that way in the same directories; they cannot be run here.

Three scripts are left alone and reported instead, because a blind edit could not
be verified: ports_module/cortex_m3 and cortex_m4 have Windows-only .bat scripts
that also name sources which are absent or differ in case, and
ports_smp/mips32_interaptiv_smp needs a MIPS toolchain.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-15 16:58:59 -04:00
Frédéric Desbiens 20d4b977f2 Removed generated build output and per-user IDE state, and fixed two stale link lines (#618)
* Removed generated build output and stopped two scripts naming deleted archives

Three kinds of file in the tree are produced by a build rather than written by
hand, and one pair of scripts still links archives that were deleted years ago.

Keil writes ThreadX_Library.plg on every build; the two committed copies are
HTML build logs from someone's machine. Code Composer generates the makefiles
under ports/c667x/ccs/example_build/tx/Release from the project files beside
them, so makefile, objects.mk, sources.mk, subdir_rules.mk, subdir_vars.mk and
ccsObjs.opt are all regenerated output.

The arm9 and arm11 sample builds link libc.a, libgcc.a and, for arm11,
libnosys.a from their own example_build directories. Those archives were removed
in 6.1.10 under "Removal of unneeded files", and libnosys.a in #594, but the link
lines were never updated, so both examples fail immediately with

    arm-none-eabi-ld: cannot find libc.a: No such file or directory

Link through the compiler driver instead, the shape every other example in the
tree uses since #594: the driver supplies libc and libgcc, and SYSCALL_LIB is
already defined in both scripts.

That does not make either example link, and the fix stops short of that on
purpose. With the archives no longer named, both now fail on

    undefined reference to `_fini'

because their linker scripts define the .init and .fini sections but not the
_init and _fini symbols, which live in crti.o and crtn.o and are omitted by
-nostartfiles. Making those two old cores build is a separate question from
removing a stale reference, so they stay in EXAMPLES_EXPECTED_TO_FAIL, with the
comment there corrected: it blamed newlib multilib packaging, which is true of
cortex_r4 and cortex_r5 but was never the reason for arm9 and arm11.

Reproduced throughout with arm-none-eabi-gcc 13.2.1.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Removed the Keil per-user state files and ignored them

Every Keil project in the tree carried a second file holding per-user state:
25 .uvoptx beside the 25 .uvprojx, 9 .uvopt beside the 9 .uvproj, and 3 .uvgui
multi-project workspace files. uVision rewrites all of them whenever a project is
opened, so they record whoever last had it open rather than anything about the
port: debugger selection, breakpoints, watch windows, window geometry.

Nothing in the tree references them, and every affected directory keeps its
.uvprojx or .uvproj, which is the file that actually describes the project.

Add ignore rules so they do not come back the next time someone opens a project
and commits.

1.4 MB across 37 files.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-15 16:40:19 -04:00
Frédéric Desbiens b1a48824ea Enabled ATCM on S32Z280, and left the other two banks off for a measured reason (#617)
entry.S programs ATCM to 0x30000000 and enables it at both exception levels. This
happens at EL2 deliberately: writing ENABLEEL2 from EL1 is silently ignored, which
was measured on BTCM and CTCM -- the base took and ENABLEEL10 took while
ENABLEEL2 stayed clear -- and the same write from EL2 sticks. tcm_enable() keeps
ENABLEEL10 as its success criterion for the same reason, since requiring both
would report failure for a bank that is usable at the level the caller runs at.

Programming ATCM also moves it off address 0, where CFGTCMBOOTx leaves it, so a
null-pointer write now faults instead of quietly landing in tightly-coupled
memory.

The boot image verifies rather than programs, preloads ATCM because ECC is enabled
and the check bits are not initialised by the core, and proves the bank holds data
both before the MPU is enabled and after. One MPU region covers it: the TRM
requires a region before an enabled TCM can be used, and an enabled TCM always
behaves as Non-cacheable Non-shareable Normal memory whatever the region says, so
only the permissions there matter.

BTCM and CTCM are left disabled, and that is a measurement rather than caution.
Enabling any second bank removes all measurable data-cache benefit:

    ATCM only                      cache gain 24%
    ATCM + BTCM                    cache gain  0%
    ATCM + BTCM + CTCM             cache gain  0%
    ATCM + BTCM at another base    cache gain  0%
    ATCM + CTCM, BTCM disabled     cache gain  0%

Five configurations, one variable. Not a particular bank, not its address, and not
the ECC preload: enabling a second bank at all. The benchmark buffer is in
non-RTU-local SRAM at 0x31800000, outside every TCM window, and CCSIDR reports the
same 16KB four-way cache throughout. I was wrong twice while narrowing this --
first blaming the preload, then blaming BTCM specifically -- and each was settled
by a run rather than by argument.

No erratum covers it. Checked the Cortex-R52 errata notice SDEN-857344 issue 19,
all twenty-five entries, and the S32Z2 0P91J mask set errata, whose RTU and R52
entries are ERR050509, ERR051107, ERR051153, ERR051441, ERR051613, ERR051614 and
ERR052126. A RAM pool shared between the RTU's last-level cache and the TCMs would
explain it, the LLC being documented as allocating ways to specific domains, but
that is a guess and it belongs with the other questions for NXP.

Little is lost meanwhile. The reference manual describes TCM_A as the bank
"optimized for small, regularly executed code such as interrupt service routines
or OS kernels", which is what a TCM is wanted for here, and enabling the other two
is one line each in entry.S once there is an answer.

Two checks were also wrong and are fixed. T5 reported every bank accessible while
two were disabled, because a disabled TCM's address range is serviced through AXIM
and memory answering there says nothing about the TCM; it now requires the bank to
be enabled as well. And T2 called tcm_enable() from EL1 for all three banks, which
would have switched on the very banks entry.S leaves off.

Verified on S32Z280 silicon: ATCM enabled and holding data before and after the
MPU, the cache benchmark back to 24%, protection checks X2 and X4 unchanged, and
the ThreadX demo still reporting 100 ticks with 20 preemptions.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-14 16:14:24 -04:00
Frédéric Desbiens ba5e5897f4 Made the basic processing counter visible to the reporting thread (#616)
The thread-metric basic processing test reports 0 for every period when built
with optimisation, and additionally claims "Basic processing thread died!"
because the counter never differs from the previous reading.

tm_basic_processing_counter is written by the processing thread and read by the
reporting thread, but it was a plain global. The processing loop calls nothing,
so nothing forces the compiler to write the counter back to memory, and in an
infinite loop there is no exit path where it must. At -O2 with
arm-none-eabi-gcc 13.2.1 the counter is loaded once before the loop, incremented
in a register, and never stored:

    28:  ldr  ip, [r3]        counter loaded once
    ...                       inner loop over the volatile array
    50:  add  ip, ip, #1      increment stays in the register
    54:  b    2c              and around again

The reporting thread reads the memory location, which stays 0 for the life of
the program. This is not specific to the Armv8-R target it was reported on: the
same shape appears for Cortex-M4 in Thumb state, and at -O1, -O2, -O3 and -Os.
Only -O0 happens to work.

Declare the counter volatile so the increment becomes a real store. Read it into
a local once per pass and use the local inside the 1024-iteration loop, rather
than letting every iteration re-read the volatile: that would add a memory access
to each iteration and change the amount of work the test performs. This test is
the baseline the rest of the suite is scaled against, per the readme, so its
throughput has to stay comparable with previously published figures and with
other RTOSes.

The array was already volatile, which is why the arithmetic itself survives
optimisation; only the counter was missing.

Verified by disassembly rather than by inspection. After the change the inner
loop is instruction-for-instruction identical to what dev generates today,
ldr/ldr/add/eor/str/add/cmp/bne, with one volatile read hoisted above the loop
and one store below it:

    30:  ldr  ip, [lr]        one read per pass, outside the loop
    34:  ...                  inner loop unchanged, back edge targets 34
    54:  add  ip, ip, #1
    58:  str  ip, [lr]        the counter is now published
    5c:  b    2c

Confirmed for cortex-r52 in Arm state and cortex-m4 in Thumb state.

Reported by @hotislandn in #480, which identified the cause correctly.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-14 14:45:55 -04:00
Frédéric Desbiens 27891f3b84 Read the TCM configuration off the part, and corrected what we thought we knew (#615)
The BSP has recorded since bring-up that the TCMs are inaccessible at reset
"because nothing has programmed the TCM region registers yet". Reading those
registers shows the reason was wrong.

    ATCM  0x0000011F   64KB, 1 wait state, ENABLED at EL2 and EL1/EL0
    BTCM  0x00000014   16KB, 0 wait states, disabled
    CTCM  0x00000114   16KB, 1 wait state, disabled

Every BASEADDRESS field is zero, so ATCM is live as 64KB at address 0x00000000,
not at the 0x30000000 the reference manual documents. The reads that faulted were
of an address the TCM is not at. ATCM's enables reset set because CFGTCMBOOTx is
tied high on this part, which the Cortex-R52 TRM gives as the one exception to
"at reset all bits are 0 apart from SIZE and WAITSTATES". BTCM and CTCM really
are disabled.

The sizes and wait states match the S32Z2 reference manual exactly -- TCMA 64KB
with one wait state, TCMB 16KB with none, TCMC 16KB with one -- so the TRM's field
layout, NXP's documented configuration and the silicon all agree. That agreement
is the point of reading before writing.

ECC is implemented and enabled: IMP_MEMPROTCTLR reads 0x00000011, both RAMPROTIMP
and RAMPROTEN set. TRM 6.2.2 therefore applies rather than being hypothetical: a
TCM location must be written before it is read, or the read reports an error --
which looks exactly like "the TCM is not accessible" and sends the reader back to
region registers that were already correct. The preload widths differ too, ATCM
needing 64-bit aligned STRD or STM where BTCM and CTCM accept 32-bit stores, so a
C loop over unsigned int would leave ATCM's check bits invalid.

tcm.c reads and decodes only; nothing is programmed here. The layouts in tcm.h are
quoted from TRM r1p3 section 3.3.94 table 3-136 and section 3.3.76 table 3-114,
not inferred from a neighbouring register: BASEADDRESS is [31:13] where
IMP_PERIPHPREGIONR uses [31:12], and assuming the analogy would have been wrong by
one bit in the same way the PRBAR shift was.

The boot image reports all of it and flags any size that disagrees with the
reference manual, so a part configured differently says so rather than being
silently assumed to match this one.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-14 14:18:50 -04:00
Frédéric Desbiens ea91beae42 Verified nested FIQ handling on S32Z280 silicon (#614)
#613 exercised the FIQ nesting routines on the FVP. What the model could not show
is whether a real GIC-600 routes Group 0 to FIQ the same way, which is the reason
to run it here. It does, and the counts match the model exactly.

The Group 0 support ports across unchanged: IGRPEN0 and BPR0 on the CPU interface,
Group 0 in the distributor, gicv3_enable_sgi_group0, gicv3_send_sgi_group0 through
ICC_SGI0R, and the separate Group 0 acknowledge and end-of-interrupt pair. entry.S
routes the EL1 FIQ vector into _tx_thread_fiq_context_save with the acknowledge
before nesting starts, and leaves FIQ unmasked on the drop to EL1 when FIQ support
is compiled in, for the same reason as on the FVP: tx_thread_stack_build only
clears a thread's F bit in that configuration.

One structural difference from the FVP cost a link. This example reports faults
through FAULT_TAIL rather than FAULT_REPORT, so the vector table needed a new
el1_fiq_entry label that falls back to fault_el1_fiq. Placing that label inside the
TX_R52_USE_THREADX_IRQ guard broke s32z280_boot.elf, which does not define it: the
vector reference is unconditional, so the label has to be too. It now sits outside
the guard and carries its own, the same shape the demo_m2 link break in #613
forced on the FVP side.

Verified on S32Z280 silicon:

    F1 FIQ delivered and dispatched            PASS
    low-priority FIQ count  = 0x00000015       21
    high-priority FIQ count = 0x00000014       20
    nested FIQ count        = 0x00000014       20 of 20 nested
    max FIQ depth           = 0x00000002
    FIQ depth now           = 0x00000000
    F2 FIQ nested inside an FIQ handler        PASS
    F3 FIQ nesting unwound to depth zero       PASS
    F4 IRQ tick undisturbed by FIQ work        PASS
    F5 lower-priority thread still scheduled   PASS
    F6 no unexpected Group 0 INTID             PASS

No regression, both configurations checked on the board. In the FIQ build the
ThreadX demo still reports 100 ticks with 20 preemptions and the boot image still
passes its cache and protection checks. In the default build the demo is unchanged
and the image links no FIQ or Group 0 symbol at all, so the work is absent rather
than dormant where it is not wanted -- worth confirming on hardware rather than
reasoning about, because the default build now reaches the FIQ vector through a new
label even though that label only branches to the fault reporter.

entry.S assembles in all four combinations of TX_R52_USE_THREADX_IRQ and FIQ
support, on both toolchains, and every file builds with GNU without warnings and
with Arm Toolchain for Embedded 22.1.0.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-14 13:34:34 -04:00
Frédéric Desbiens 19d90a49c4 Exercised the nested FIQ path, the last pair nothing had ever called (#613)
_tx_thread_fiq_nesting_start and _tx_thread_fiq_nesting_end complete the set:
after #611 and #612 covered IRQ nesting on the model and on silicon, these two
were the remaining routines compiled into every build and entered by nothing.
demo_fiq.elf enters them.

FIQ needs more of the GIC than IRQ does. With a single security state the
controller delivers Group 0 as FIQ and Group 1 as IRQ, so an interrupt only
arrives as an FIQ if it has been moved into Group 0, the distributor and the CPU
interface both have Group 0 enabled, and it is acknowledged through the Group 0
registers. Group 1's acknowledge returns the spurious INTID for a Group 0
interrupt and leaves it pending, which would present as a storm rather than as an
error. gicv3.c gains IGRPEN0, BPR0, gicv3_enable_sgi_group0,
gicv3_send_sgi_group0 and the Group 0 acknowledge and EOI pair. ICC_SGI0R differs
from ICC_SGI1R only in opc1, 2 against 0, and each raises into its own group.

entry.S routes the EL1 FIQ vector into _tx_thread_fiq_context_save with the same
ordering the IRQ path needed: acknowledge in FIQ mode before nesting starts, then
nesting_start, service, nesting_end, and end-of-interrupt last. It also leaves
FIQ unmasked on the drop to EL1 when FIQ support is compiled in, because
tx_thread_stack_build only clears a thread's F bit in that configuration and
nothing else ever clears it, so an FIQ raised before the first thread ran would
otherwise be silently ignored.

Nesting an FIQ means taking an FIQ while an FIQ handler runs, which one source
cannot show, so two Group 0 SGIs are used with the second at a numerically lower
priority. The low one's handler raises the high one.

One mistake worth recording, because the guard it needed is not obvious.
TX_ENABLE_FIQ_SUPPORT is PUBLIC on the threadx target, so it reaches every image
as soon as the library is built with FIQ -- including images that link no
interrupt controller at all. demo_m2 is one of those: no gicv3.c, no
irq_dispatch.c. Referencing gicv3_acknowledge_group0 and board_fiq_service from
the FIQ vector broke its link outright. The vector is now gated on
TX_R52_USE_THREADX_IRQ as well, which is how the IRQ vector has always been
gated, and images without the infrastructure keep the fault reporter.

Verified on FVP_BaseR_AEMv8R. F1 covers the whole Group 0 chain on its own --
IGRPEN0, the ICC_SGI0R encoding, IAR0 and EOIR0, the EL1 vector and F being
unmasked -- so a failure there points somewhere other than the nesting routines.
The demo reports 21 low-priority FIQs, 20 high-priority, 20 of them nested, max
depth 2, depth unwound to zero, the IRQ tick undisturbed, threads still
scheduled and no unexpected Group 0 INTID.

All six images pass in the FIQ configuration, and the default build links no FIQ
or Group 0 symbol at all, so the work is absent rather than dormant where it is
not wanted. Both toolchains build every file, GNU with no warnings and Arm
Toolchain for Embedded 22.1.0 as well.

Not done here: the S32Z280 example is untouched, so FIQ on silicon is a separate
change, as IRQ nesting was.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-14 11:42:59 -04:00
Frédéric Desbiens 5a7f4f8f7c Verified nested IRQ handling on S32Z280 silicon (#612)
#611 exercised the nesting routines on the FVP and established what breaks them:
the interrupt must be acknowledged before nesting starts, or the still-pending
level-asserted timer is retaken the moment IRQ is enabled and recurses until the
stacks are gone. entry.S here carries the same ordering for the same reason, and
irq_dispatch.c splits the same way, board_irq_service taking an
already-acknowledged INTID while board_irq_handler keeps its old shape as the
non-nesting entry point.

What the model could not answer is whether a real GIC-600 agrees, and two things
could have differed.

The first is the number of implemented priority bits. Equal priorities do not
preempt and it is the low bits that vanish, so if the timer and the SGI collapse
to one value after truncation then nesting cannot happen at all -- and the test
would fail without saying why. gicv3_priority_bits discovers the count by writing
0xFF to a priority byte and reading back which bits stick, board_init records the
two effective values, and check P1 requires the SGI to still outrank the timer.
This silicon keeps five bits, the same as the FVP, so 0xA0 and 0x50 stay distinct;
that is now measured and reported rather than assumed.

The second is whether an SGI raised on real hardware is delivered at all. ICC_SGI1R
is a 64-bit AArch32 CP15 register whose encoding does not transcribe from the
AArch64 alias, so check N1 raises one from thread context and requires delivery
before nesting is involved. It arrives.

Verified on S32Z280 silicon:

    priority bits   = 0x00000005
    timer effective = 0x000000A0    sgi effective = 0x00000050
    P1 SGI outranks timer after truncation     PASS
    N1 SGI delivered and dispatched            PASS
    max depth = 0x00000002   nested SGIs = 0x00000032   depth now = 0x00000000
    N2 SGI nested inside another handler       PASS
    N3 nesting unwound to depth zero           PASS
    N4 tick still advancing after nesting      PASS
    N5 lower-priority thread still scheduled   PASS
    N6 no spurious or unexpected interrupts    PASS

Fifty nested SGIs across fifty ticks, one per tick, and depth never exceeded two.

No regression, checked in both configurations on the board. In the nesting build
the ThreadX demo still reports 100 ticks with 20 preemptions and the boot image
still passes its cache and protection checks. In the default build the image links
no nesting symbols at all and the demo is unchanged. Both toolchains compile every
file, GNU with no warnings and Arm Toolchain for Embedded 22.1.0 as well.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-14 11:07:06 -04:00
Frédéric Desbiens 559460bed9 Exercised the nested IRQ path, which had never once been entered (#611)
_tx_thread_irq_nesting_start and _tx_thread_irq_nesting_end have shipped in this
port since it was written, compiled into every build, and nothing had ever called
either one. Not on the model, not on silicon, not in any demo. demo_nesting.elf
enters them.

Provoking nesting needs two sources with different priorities. The generic timer
PPI was already there; the second is an SGI, which a core can raise on itself.
gicv3.c gains gicv3_enable_sgi and gicv3_send_sgi for that. ICC_SGI1R is 64-bit,
so in AArch32 it is an MCRR rather than an MCR, and the AArch64 name the
Cortex-A72 example uses, S3_0_C12_C11_5, does not transcribe to the AArch32 CP15
space. The encoding here is confirmed by check N1 in the demo, which raises an SGI
from thread context and requires it to be delivered and dispatched.

The order of the pairing is the whole difficulty, and getting it wrong does not
fail gracefully. The interrupt must be acknowledged BEFORE nesting starts. Reading
ICC_IAR1 is what raises the GIC running priority to this interrupt's own, which
masks it and everything of equal or lower priority; only then is re-enabling IRQ
safe. My first attempt called nesting_start first and acknowledged inside the
handler, so the still-pending, still-level-asserted timer was taken again the
instant IRQ was enabled, and again, until the IRQ and System stacks were
destroyed. It presented as garbage on the console and a hang with no fault to
point at, and it broke demo_m3 and demo_threadx while leaving boot_check, demo_m2
and demo_mpu passing, because only the first two depend on the tick advancing. The
Cortex-R5 example BSP states the requirement in one line: "ensure all IRQ
interrupts are cleared prior to enabling nested IRQ interrupts."

So entry.S now acknowledges in IRQ mode, carries the INTID in r4 -- which survives
the mode switch, since only SP and LR are banked -- and also pushes it on the IRQ
stack so a nested level reusing r4 cannot lose the outer level's value.
End-of-interrupt waits until after nesting_end, in IRQ mode with interrupts
masked, so dropping the running priority cannot re-admit the same interrupt.

board_irq_handler splits in two. board_irq_service does the middle part on an
already-acknowledged INTID and neither acknowledges nor EOIs; board_irq_handler
keeps its old shape as the non-nesting entry point, so images built without
TX_ENABLE_IRQ_NESTING behave exactly as before.

The nesting instrumentation in irq_dispatch.c is inert unless an image asks for
it. board_nest_provoke gates the SGI that the timer handler raises, so every other
image sees the handler it always had.

Verified on FVP_BaseR_AEMv8R. The nesting demo reports max depth 2, fifty nested
SGIs across fifty ticks, depth unwound to zero, the tick still advancing, the
low-priority thread still scheduled, and no spurious or unexpected INTIDs. In the
same nesting configuration the five existing images all pass, so the split did not
disturb the ordinary path. Both toolchains build it, GNU with no warnings and Arm
Toolchain for Embedded 22.1.0 as well.

Not done here: the S32Z280 example keeps its own irq_dispatch.c and entry.S and is
untouched, so nesting on silicon is a separate change.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-13 17:13:34 -04:00
Frédéric Desbiens b877305d13 Verified lazy VFP context switching on S32Z280 silicon (#609)
The FVP proved the lazy save and restore path at AR1/M5, but the board had never
run it. Doing so needed one thing the model did not: the FPU has to be turned on.

entry.S opens CPACR for CP10/CP11 and sets FPEXC.EN at EL1, guarded by __ARM_FP.
EL2 already cleared HCPTR.TCP10/TCP11, but both of the EL1 gates read 0 out of
reset, so a floating-point instruction raised an Undefined Instruction exception
before this. tx_thread_vfp_enable() does not help: it sets the per-thread software
flag that makes the context switch save and restore the registers, and never
touches the hardware. Enabling it is the BSP's job. The __ARM_FP guard is what
keeps a soft-float build assemblable, since "vmsr fpexc" is not a valid
instruction for that target at all.

demo_vfp_s32z280.c is the FVP's demo_m5.c test design over the LINFlexD console.
The design is kept deliberately, because the two halves of the VFP context path
need separate provocation: "fp check" holds eight live doubles across
tx_thread_sleep, eight being enough to force the callee-saved D8-D15 bank that a
solicited switch must preserve, while "fp busy" sits at the lowest priority and
never sleeps, so the tick interrupts it mid-computation and exercises the
interrupt half, D0-D15 plus FPSCR. Only these two threads opt in, so the test
also shows that opting in is what does the work.

s32z280_vfp.elf is gated on TX_R52_ENABLE_VFP, as demo_m5.elf is in the FVP
example, since the image is meaningless unless the library was built with a
floating-point ABI.

Also made this example's -Wl,--no-warn-rwx-segments conditional on the compiler
being GNU. #604 did that for the FVP example and this file still carried the
literal, so ld.lld failed the link with "unknown argument". With that fixed the
S32Z280 images build with Arm Toolchain for Embedded too.

Verified on S32Z280 silicon:

    V2 D8-D15 bank preserved across switches   PASS   50 solicited switches
    iterations = 0x000140A8                           82,088 interrupted rounds
    corruptions = 0x00000000
    V3 interrupted FP thread made progress     PASS
    V4 no FP corruption across interrupts      PASS
    filex_ptr = 0xF11EF11E
    V5 VFP flag did not alias filex_ptr        PASS
    PASS lazy VFP context switch verified on silicon

No regression: the soft-float build is warning-free and its image contains no vmsr
at all, confirming the guard elides the block, and the existing demo still reports
100 ticks with 20 preemptions on the board. Both S32Z280 images also build with
Arm Toolchain for Embedded 22.1.0, the hard-float VFP image included. No FVP file
is touched.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-13 17:07:36 -04:00
Frédéric Desbiens c991c9e6ae Added TX_ENABLE_FIQ_SUPPORT to the feature-macro assembly stage (#610)
#608 assembled the code behind TX_ENABLE_VFP_SUPPORT, TX_LOW_POWER and
TX_ENABLE_EXECUTION_CHANGE_NOTIFY, and missed TX_ENABLE_FIQ_SUPPORT, which guards
assembly in 145 files across the A and R profile ports. All 145 assemble today, so
this adds no fix, only the regression protection the other three already have.

Also recorded why TX_ENABLE_IRQ_NESTING and TX_ENABLE_FIQ_NESTING are not in the
list, since their absence otherwise looks like the same oversight. They guard no
assembly in the trees this script walks: the nesting start and end routines are
separate files compiled unconditionally, and the macros only feed the
TX_PORT_SPECIFIC_BUILD_OPTIONS bitfield in tx_port.h. Adding them would assemble
nothing new while implying coverage that does not exist.

Verified with Arm Toolchain for Embedded 22.1.0: 711 of 711 assembly sources, then
37 of 37 VFP, 145 of 145 FIQ, 8 of 8 TX_LOW_POWER and 218 of 218
TX_ENABLE_EXECUTION_CHANGE_NOTIFY.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-13 15:57:18 -04:00
Frédéric Desbiens acdc02b2fc Assembled the code behind feature macros, and fixed the POP it found (#608)
scripts/check_clang.sh assembled every source with default flags, so the
preprocessor discarded each #ifdef block before the assembler saw it. Nothing in
the tree had ever assembled a guarded path. That covers the VFP context save and
restore in ten ports, and 218 files carrying TX_LOW_POWER or
TX_ENABLE_EXECUTION_CHANGE_NOTIFY.

Turning those on found a defect. The Cortex-M0 and Cortex-M23 execution-profile
paths bracket their call with

    PUSH    {r0, lr}
    BL      _tx_execution_isr_enter
    POP     {r0, lr}

and the last of those is invalid on Armv6-M and Armv8-M Baseline, where the
16-bit Thumb POP takes r0-r7 and pc and nothing else. GNU rejects it as well --
"cannot honor width suffix" -- so TX_ENABLE_EXECUTION_CHANGE_NOTIFY and
TX_EXECUTION_PROFILE_ENABLE have never been buildable on either port with either
toolchain. Four files, all the same shape.

The fix pops into a scratch register and moves it, MOV to a high register being
permitted where POP is not. r1 is free: the BL may clobber r0-r3, which is the
reason r0 is saved in the first place. Disassembling the result gives
push {r0, lr} / bl / pop {r0, r1} / mov lr, r1 / bx lr, one 16-bit instruction
more than before and otherwise the same.

Two findings that were not defects, recorded in the script so they are not
rediscovered:

Cortex-R4 needs an -mfpu to assemble its VFP path, because its FPU is an option
rather than part of the core. GNU fails identically without one, so this is a
flags requirement and not a toolchain divergence.

The A profile ports must not be given one. Adding -mfpu=vfpv3-d16 uniformly broke
28 files with "register expected", because those ports save D16-D31 and a -d16
FPU does not have those registers. Their defaults were already right.

The new stage runs under --asm-only as well, needing no target C library, and
reports 37 of 37 VFP files, 8 of 8 TX_LOW_POWER and 218 of 218
TX_ENABLE_EXECUTION_CHANGE_NOTIFY. Restoring the POP for one run makes it fail
with 217 of 218 and name the file and the error, so the stage is not vacuous.
The other four stages are unchanged: 711 of 711 assembled, 185 of 185 common C
sources for each of nine cores, 42 of 42 script-driven examples and 5 of 5 CMake
images.

The fixed code is verified to assemble with both toolchains and to encode as
intended. It is not verified running: there is no Cortex-M0 or Cortex-M23 model
here, and these are context save and restore paths, so that gap is worth stating.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-13 10:14:53 -04:00
Frédéric Desbiens 010a6c9fbb Gave the gnu ports a CMake build, which most of them lacked (#607)
The project guidelines ask for CMake and Ninja, but only 15 of the 59 gnu port
directories had a CMakeLists.txt. None of the 27 AArch64 ports had one, so the
architecture whose examples were repaired over the last few changes still could
not be built the way the project says to build it, and nothing in CI could
compile it.

Add a CMakeLists.txt to the 44 that lacked one. Three of them are templates in
ports_arch, because 34 of the 44 are generated: the ARMv7-A and AArch64 source
lists are uniform within each family, so one template per family serves every
core in it and update.sh distributes it. The other 10 ports have no generator
and get their own file.

Add the toolchain files those ports select, following the shape of
cmake/cortex_a9.cmake. AArch64 needs a base file of its own rather than a
variant of arm-none-eabi.cmake: it has no -marm or -mthumb to choose between and
no -mfloat-abi, and aarch64-none-elf-gcc rejects -mlong-calls outright, so that
flag cannot be carried across. The tools are named without a path, unlike
cmake/cortex_r52.cmake which pins one, because pinning 30 files to a single
machine's directory layout is the problem the previous change removed from the
launch configurations.

Three toolchain files cover ports that already had a CMakeLists.txt but no way
to select it: the Armv8-M mainline gnu ports, cortex_m33, cortex_m55 and
cortex_m85. Without cmake/<arch>.cmake the documented invocation cannot reach
them.

The top level needed one fix. It derives the SMP port directory as
<arch>_smp, but ports_smp/linux and ports_smp/win64 predate that convention and
carry no suffix, so those two could never be configured. Fall back to the bare
name when the suffixed directory is absent. The check only fires when the
suffixed directory does not exist, so no port that already resolved changes
behaviour, and ports_smp/win64's existing CMakeLists.txt becomes reachable too.

Verified by configuring and building every one: 53 of 53 static libraries build
with cmake -G Ninja, using Arm GNU Toolchain 14.3.Rel1 for both arm-none-eabi
and aarch64-none-elf. That covers the 44 new ports plus the 9 that already
worked, and includes ports_smp/linux, which failed before the fallback.
scripts/check_ports.sh passes, so the three templates and their 34 generated
copies agree.

Six gnu ports are still outside the CMake build, all for want of a compiler
rather than a CMakeLists.txt: rxv1, rxv2 and rxv3 need the Renesas RX GNU
toolchain and mips32_interaptiv_smp needs a MIPS one, neither of which is
available here, so writing toolchain files for them would mean shipping
untested guesses. risc-v32 and risc-v64 already build through their own
differently named toolchain files.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-12 12:43:57 -04:00
Frédéric Desbiens 990d51b65d Brought up ThreadX on S32Z280 silicon and fixed the MPU encoding it exposed (#606)
* Added an S32Z280-594EVB example build for the Cortex-R52 port

First silicon bring-up for the Cortex-R52 port: an image that boots RTU0 core 0
on the NXP S32Z280-594EVB, drops from EL2 to EL1 and reports what the core says
about itself.  Gated behind TX_R52_BUILD_S32Z280_EXAMPLE, separate from the FVP
example option so that FVP regression is unaffected by bring-up work and the
two boards' differing reset states cannot interact.

Verified on the board.  The image runs from reset and reaches its completion
breakpoint, and the two CPSR values are the point of the exercise: mode 0x1A
(Hyp) at EL2 then 0x13 (Supervisor) at EL1, i.e. the EL2 to EL1 drop performed
by this code on silicon rather than on a model.

    MIDR      0x411FD133   Cortex-R52 r1p3, part 0xD13
    MPUIR     0x00001400   20 MPU regions at EL1
    HMPUIR    0x00000014   EL2
    CTR       0x8144C004
    MPIDR     0x80000000
    ID_PFR0   0x00000131
    ID_PFR1   0x10111001   virtualisation field 1: EL2 implemented
    SCTLR     0x70C50838   MPU, D-cache and I-cache all off at reset
    CNTFRQ    0x00000000

The FVP reports MIDR 0x410FD0F0, part 0xD0F, which is an architecture envelope
model and not a Cortex-R52 -- so these are the first implementation-specific
numbers the port has had.

Three silicon facts the code exists to encode, each of which presents as
working hardware that quietly does the wrong thing:

  - The core resets in THUMB state.  CPSR reads 0x1FA out of reset because the
    boot instruction NXP plants at the boot address is a T32 branch, so _start
    is T32 and switches to A32 itself.  An A32 entry would execute the first
    halfword of its own instruction as Thumb.

  - The debugger holds every core in debug state.  NXP's
    _reset_to_first_instruction() asserts MDM_AP CONTROL2[19:16] =
    CR52_RTU0_{3,2,1,0}_EDBGREQ and never clears them, though its own comment
    says start_debug_by_core_name() does.  While asserted the core executes
    nothing, yet registers and memory still respond and MC_ME/RGM report the
    core released and clocked.  tools/read_identity.gdb clears core 0's bit and
    verifies the clear took.

  - CNTFRQ reads zero, exactly as on the FVP, and is writable only at the
    highest implemented exception level.  entry.S records it rather than
    writing a value, because the correct frequency for this board is not yet
    established and a wrong one would silently mis-scale every derived
    interval.  Timer work must program it first.

Memory map: .text is linked and loaded at 0x79900000, the instruction-fetch
window and the reset address (MC_ME_PRTN0_CORE0_ADDR reads exactly that).  It
is writable over the debug AXI port despite NXP's map declaring it read-only,
so no load-address alias is needed for a debugger-loaded image; a flash-booted
one would use the data alias at 0x32100000, which is the same physical SRAM as
confirmed by a sentinel write.  Data, bss and the per-mode stacks go to
RTU-local data SRAM at 0x31780000.  The TCMs are absent on purpose: they are
inaccessible at reset until their region registers are programmed, so nothing
needed for early boot can live there.

Reporting is built so that a failure cannot read as a success: boot_stage
distinguishes a fault from a hang and fault handlers record vector, syndrome
and address before parking; r52_identity.magic is written last so a
half-filled structure is detectable; and bsp_done() is a deliberate breakpoint
target so completion is observed rather than inferred from a timeout.

The image does not link ThreadX, so a failure here is unambiguously a boot or
board problem rather than a kernel one -- the same reasoning as the FVP's
boot_check.elf.

Not included: console, timer, GIC, MPU programming.  LIN9 reaches the host
through the daughtercard USB-UART, which is LINFlex_9 at 0x42980000 and not
LINFlex_0; what is missing is the clock configuration needed to compute a baud
rate, and a console at the wrong rate produces garbage indistinguishable from a
crash.  Hence reporting through memory for now.

FVP example builds and passes 6/6 unchanged.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Added a LINFlexD console to the S32Z280-594EVB example

Polled transmit console on LINFlex_9, which is the instance wired to the
daughtercard USB-UART through jumper J248.  Not LINFlex_0: that is merely the
first instance in the Reference Manual's list and reaches no connector on this
board.  bsp_boot.c now prints the identity registers as well as recording them
in memory; the memory copy stays, because it is what proved the boot path
before a console existed and it still works if the console is misconfigured.

The baud rate was derived rather than assumed, which mattered because a console
at the wrong rate produces garbage indistinguishable from a crash:

  - LINFlex_9 is clocked by P5_LIN_BAUD_CLK, driven by MC_CGM_5 MUX 2.
  - Read from the board: MUX_2_CSS selects source 2, MUX_2_DC_0 divides by 1.
  - The EVB carries a 40 MHz crystal (UG10268 section 3.3.1.1), so 40 MHz.
  - LFDIV = 40000000 / (16 * 115200) = 21.7014, giving LINIBRR 21 and
    LINFBRR 11, for an actual 115274 baud -- 0.06% error.

Those are exactly the values the BootROM had already left in the registers,
which is the corroboration for the 40 MHz figure rather than merely a
plausible-looking calculation.

They are programmed here anyway instead of inherited, for two reasons.
Inheriting register state makes this image's behaviour depend on how the board
was last booted.  And the BootROM's own configuration is wrong for a console:
it leaves PCE set and TxEn clear, because it is listening for a serial-boot
download rather than printing, so a host terminal on 8N1 would see framing
errors.

Verified on the board: the image reaches its completion breakpoint with the
console calls in the path, so every byte was accepted and UARTSR.DTF was set
for each -- the transmitter is configured, enabled and clocked.  That does not
by itself prove the rate, since DTF sets at any baud; the rate rests on the
arithmetic above and the BootROM's independently matching divisors.  Observing
the text needs the daughtercard USB-UART connected to a host terminal at
115200 8N1.

Register offsets and bit positions are from Reference Manual section 75.5.1
and the UARTCR/UARTSR diagrams.  PCE at bit 2 and TxEn at bit 4 are called out
in the source because they are easy to transpose and the failure is silent.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Fixed two silent failures in the S32Z280 LINFlexD console

The console transmitted from the first attempt, but what arrived on the wire
was wrong in two independent ways, neither of which the target could detect.
Both are now fixed and the output is byte-for-byte correct, captured from the
USB-UART rather than read off a terminal.

1.  UARTCR was written outside initialisation mode.

    Setting LINCR1.INIT does not put the module in initialisation mode in the
    same cycle, and most of UARTCR is writable only there.  Configuring
    immediately after the LINCR1 write half-worked: bits being *set* took
    effect while bits being *cleared* did not.  TxEn came on, but PCE stayed
    on, so the line ran 8E1 against a host expecting 8N1.  That corrupts only
    those characters whose parity bit happens to be 0 and leaves the rest
    readable, which looks like a marginal baud rate rather than a framing
    error -- and would have sent the next hour into the clock tree.

    linflexd_init() now polls LINSR[LINS] until the module reports
    initialisation mode, reads UARTCR back afterwards, and returns a status
    mask.  bsp_boot.c records it and prints it as CONSOLE; 0 means both the
    mode entry and every field write were confirmed rather than assumed.

2.  DTF was cleared after the byte rather than waited on.

    DTF is write-one-to-clear and does not de-assert in the same cycle as the
    clearing write, so the next byte's poll could observe the previous byte's
    flag, conclude the line was free while it was still busy, and have its own
    write silently discarded.  That cost exactly one character after every
    "\r\n" pair -- the only place two bytes go out back to back -- so the
    first letter of every line went missing: MIDR read as IDR, SCTLR as CTLR.

    The transmit sequence is now write, wait for DTF, clear DTF, then wait for
    the clear to take effect, with the same bounded guard as the init loop so
    a stuck flag degrades to slow output rather than a hung boot.

    Clearing before the write was tried and is worse, not better: the write
    then lands while the previous byte is still shifting and is dropped, and
    the poll afterwards sees the previous byte's completion, so most of the
    output disappears.  That failure is what established that the transmitter
    discards writes while busy, which is the fact both bugs turn on.  The
    source says so, because the ordering looks arbitrary otherwise.

Verified end to end on the board, with the USB-UART passed through to WSL2 so
the received bytes could be compared against what the code intended to send:

    === ThreadX Cortex-R52 :: NXP S32Z280-594EVB ===
    MIDR      = 0x411FD133
    MPUIR     = 0x00001400
    HMPUIR    = 0x00000014
    CTR       = 0x8144C004
    MPIDR     = 0x80000000
    ID_PFR0   = 0x00000131
    ID_PFR1   = 0x10111001
    SCTLR     = 0x70C50838
    CNTFRQ    = 0x00000000
    CPSR@EL2  = 0x000001DA
    CPSR@EL1  = 0x600001D3
    EL1 MPU regions: 00000014
    CONSOLE   = 0x00000000
    === boot complete ===

Reaching the completion breakpoint proved only that the transmitter ran; it
could not have caught either fault.  Comparing bytes is what did.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Added generic timer support to the S32Z280-594EVB example

The Arm generic timer now runs on silicon: CNTFRQ is programmed at EL2 and a
one-shot compare fires when it should.

    CNTFRQ programmed = 0x007A1200 (8000000)
    elapsed counts    = 8000043   (requested 8000000, +0.001%)
    host interval     = 1.0033 s

Both halves matter.  The elapsed count shows the counter and the compare agree
with each other; the host interval shows the rate is actually 8 MHz rather than
merely self-consistent, which a wrong CNTFRQ would also produce.  The residual
43 counts is about 5 us of polling overhead.

CNTFRQ reads zero out of reset here, exactly as on the Armv8-R AEM FVP -- so
that FVP behaviour is real silicon behaviour, not a model artefact, and the
port readme's warning to re-verify it was well placed.  The parallel stops
there: the FVP also leaves the system counter itself stopped and its BSP must
start it, whereas on this board the BootROM has the counter running and only
the software-declared constant was missing.  CNTFRQ is writable only at the
highest implemented exception level, so entry.S programs it before the drop to
EL1.

8 MHz was established three independent ways rather than assumed:

  - Measured: CNTPCT sampled against host wall-clock time over a 32-second
    interval gave 8.0227 MHz.
  - Derived: RTU.GPR CFG_CNTDV reads 4, so the divider is (4+1) = 5, and the
    board's FXOSC is 40 MHz -- itself already corroborated by the LINFlexD
    baud divisors the BootROM left behind.  40 / 5 = 8.
  - Confirmed: the one-shot compare above.

The compare is polled with the interrupt masked, deliberately.  There is no
GIC configured yet, and proving the timer counts and fires first means a later
interrupt failure cannot be confused with the timer itself being wrong -- the
same staging that made the console tractable.

Also recorded, from a detour that cost a run: RTU0.GPR CFG_CNTDV at 0x76120010
is readable by the debugger, which returns 4, but a load from the core at EL1
kills the image.  No fault handler runs, boot_stage stays at 4, and the debug
connection drops -- the signature of a stalled bus access rather than an abort,
and a failure mode that leaves nothing on-target to inspect.  The debugger
reaches that window over the AXI-AP, which does not go through whatever gates
the core's own access to RTU peripheral space.  The image no longer touches it;
the value lives in platform.h instead.  This is the second address on this part
that is debugger-visible but not core-reachable, so peripheral windows are now
worth probing from the core before relying on them.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Enabled the R52 peripheral port and located the GIC on S32Z280

Two peripheral windows on this part stalled the core outright when read: no
abort, no fault handler, boot_stage frozen at its last value, and the debug
connection dropping with it.  A stall gives you nothing on-target to inspect,
so both were found by printing a marker before each access and seeing which
marker came last.  The Cortex-R52 TRM r1p3 (100026_0103_00_en) explains both;
NXP's reference manual mentions neither.

1.  The low-latency peripheral port is disabled at reset.

    IMP_PERIPHPREGIONR (TRM 3.3.80) describes a region at 0x76000000 of 4 MB
    on this part -- the RTU peripheral space -- with separate enables for EL2
    and EL1/0.  Read from the board it was 0x76000034: base 0x76000000, size
    0b01101 = 4 MB, and both enable bits clear.  The TRM is explicit that each
    "resets to 0".  Until they are set, every access in that window stalls.

    entry.S now sets both at EL2, which is also where it has to happen: EL1
    writes to this register trap to EL2 when HACTLR.PERIPHPREGIONR is clear.
    PERIPHPRG now reads 0x76000037 and the core reads RTU0.GPR CFG_CNTDV = 4
    for itself -- the same value the debugger saw, which independently
    confirms the divider behind the 8 MHz counter rate from the core's own
    view rather than the debug path's.

2.  The GIC is where NXP says, and needs an MPU mapping to reach.

    IMP_CBAR (TRM 3.3.17) holds the physical base of the memory-mapped GIC
    distributor in bits [31:21], its reset value wired from CFGPERIPHBASE.
    Read from the board it is 0x47800000, confirming NXP's memory map from the
    hardware rather than from a vendor debugger script -- the reference manual
    contains no GIC base address anywhere.  NXP's cryptic note that the "R52
    Cluster set addr [31:21]" is just CFGPERIPHBASE[31:21] restated.

    TRM Table 9-1 gives the frame layout relative to that base: distributor at
    +0x000000, redistributor control at +0x100000, redistributor SGI/PPI at
    +0x110000, then +0x20000 per further core.  That matches what the FVP
    example's driver already assumes, so it should port with address changes.

    The distributor still stalls, and the TRM says why: "Ensure that the memory
    region used for the GIC Distributor is configured as Device nGnRnE."  The
    MPU is disabled here -- SCTLR.M reads 0 -- so nothing maps that region.
    Enabling the peripheral port does not help, the GIC being outside that
    window and reached over AXIM.  Reading it waits on MPU programming, which
    is the next piece of work; the probe is removed rather than left in, since
    while present it stalled every run and cost the console and timer results
    that precede it.

gic_probe.c holds the two system-register reads.  They cannot hang the bus,
unlike the memory accesses they diagnose, which is the point of them.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Reached the GIC on S32Z280 by mapping it Device nGnRnE

The GIC distributor and redistributor now respond on silicon:

    MPU rgns   = 0x00000003
    SCTLR      = 0x70C50839      (M set; caches still off)
    GICD_PIDR2 = 0x0000003B      architecture revision 3 = GICv3
    GICD_TYPER = 0x0248001E
    GICR_PIDR2 = 0x0000003B      redistributor at base + 0x100000

Three independent sources had to agree to get here, and NXP's reference manual
supplied none of them.  IMP_CBAR reports the distributor base from the hardware
(0x47800000).  Cortex-R52 TRM Table 9-1 gives the frame layout: distributor at
+0x000000, redistributor control at +0x100000, SGI/PPI at +0x110000, then
+0x20000 per further core -- the layout the FVP example's gicv3.c already
assumes.  And the TRM requires the region be Device nGnRnE, which was confirmed
rather than taken on faith: mapped Normal write-back the distributor stalls the
core outright, mapped Device it reads 0x3B.

mpu.c is ported from the FVP example, keeping its PRBAR/PRLAR encoding and the
reversed-AP-bit-order finding, with an S32Z280 region table.  Caches are left
off: the point of this pass was reaching the GIC, and the debugger writes this
image straight into SRAM behind the caches, so enabling C and I deserves its own
step and its own check.

Diagnosis needed two new pieces of machinery, both kept:

  - Memory-resident progress markers (probe_stage), mirroring the console
    markers.  Enabling the MPU can take the core down in a way that records no
    fault AND takes the console with it, so boot_stage alone could not say how
    far a risky sequence got.  With markers inside mpu_init the failure went
    from "somewhere in the MPU code" to "the SCTLR.M write itself" in one run.

  - Resolving symbol addresses from the current ELF with nm on every
    post-mortem.  Adding probe_stage to the .data block shifted fault_vector,
    so reusing addresses across builds reads the wrong words -- which nearly
    produced a conclusion from a stale read.

OPEN QUESTION, recorded in mpu.c rather than papered over.  A five-region map
that adds a second Device window for the RTU peripheral space (0x76000000, 4 MB)
makes the SCTLR.M write kill the core: all regions program without error,
probe_stage reaches 0x42, and the marker one instruction later never lands.  No
exception is taken, so it stalls rather than aborts.  Coverage is not the
explanation -- that map spanned the whole address space, and an earlier comment
of mine claiming otherwise was wrong and has been corrected.  The cause is not
known.  Consequence: RTU peripheral access after the MPU is enabled is untested;
the CFG_CNTDV read this image does happens before the enable.

Also not tightened yet: code is writable and data executable in this map.  Wrong
for a protection demo, right for bring-up where an unmapped address costs a
silent stall rather than a fault.  Narrowing it belongs with the ThreadX
integration, which can verify it.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Fixed an out-of-bounds MPU table write that made enabling the MPU stall the core

The five-region map now programs correctly and the MPU enables with the GIC
reachable:

    R  idx  PRBAR      PRLAR
    R  0    00000000   477FFFC1
    R  1    47800002   479FFFC3    GIC, XN, Device nGnRnE
    R  2    47A00000   75FFFFC1
    R  3    76000002   763FFFC3    RTU peripherals, XN, Device nGnRnE
    R  4    76400000   FFFFFFC1
    MPU rgns = 5, SCTLR = 0x70C50839 (M set)
    GICD_PIDR2 = 0x3B, GICD_TYPER = 0x0248001E, GICR_PIDR2 = 0x3B

The bug was mine and it was simple: mpu_regions[] is declared with a fixed
size, and the FVP example this file came from sized it [3] to match the three
regions it programs.  A five-region table wrote mpu_regions[3] and [4] past the
end of the array.

What made it hard to see is what the corruption produced.  Every region
appeared to program without error, and the failure was a stall on the SCTLR.M
write -- no abort, no fault handler, no exception, and the debug connection
dropping.  Reading the regions back out of the hardware, with the MPU
deliberately left disabled so the readback could not itself stall, showed
region 3 with base 0x00000000 instead of 0x76000000.  With its correct limit
that region spanned 0x00000000-0x763FFFFF and overlapped regions 0 to 2, and
PMSAv8-R leaves overlapping regions UNPREDICTABLE -- so stalling was permitted
behaviour rather than a hardware fault.

One bug accounts for every observation: three regions worked, five failed, and
it failed identically whether the extra window was Device or Normal, because
the memory type was never involved.

Fixes: the table capacity is now named (MPU_TABLE_REGIONS) and sized 16, and
mpu_init checks mpu_regions_used against it as well as against MPUIR.  The
existing check only compared against the hardware's region count, which said 20
and told us nothing about the array.

Two earlier claims of mine, corrected here rather than left standing.  A comment
asserting "narrow coverage was the problem" was wrong: the failing map covered
the whole address space.  And this is NOT a defect inherited from the FVP
example -- that code declares [3] and uses three regions, entirely
self-consistent.  Worth offering upstream only as hardening: the bare [3] with
no bound check is what let a growing map overrun it silently.

The region readback is kept in the image.  It costs eight lines a boot and
verifies the map against the hardware every time, which is precisely what
turned an unexplained stall into a one-line diagnosis.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Enabled interrupt-driven ticks by clearing SCTLR.TE

Timer interrupts are delivered, acknowledged and re-armed on silicon:

    SCTLR       = 0x30C50838 -> 0x30C50839   (TE cleared, M set)
    IRQ count   = 7
    timer INTID = 30                          measured, not assumed
    spurious    = 0
    unexpected  = 0

SCTLR.TE -- Thumb Exception enable -- RESETS SET on this part.  Every vector
table in entry.S is A32, so the core entered each exception in T32 state,
decoded an A32 branch as Thumb, landed mid-instruction at a misaligned address
and took an undefined-instruction exception, which repeated the same way.  No
handler ever ran.  entry.S now clears it at EL1 and clears HSCTLR.TE at EL2.

The consequence was worse than a crash: it made every fault invisible.
boot_stage stayed at its last value, fault_vector stayed zero, and the core sat
at el1_vectors+4.  A data abort from an unmapped peripheral, an abort from an
overlapping MPU region, and a correctly delivered timer interrupt all presented
identically -- an unexplained hang with nothing recorded.  What gave it away was
reading the banked registers: LR_irq and SPSR_irq showed an IRQ had been taken
from SVC with interrupts enabled, CPSR showed UND mode with T set, and LR_und
pointed at a misaligned address inside the A32 vector table.

The Armv8-R AEM FVP resets with TE clear, which is why the same vector tables
work there and why nothing in the FVP work anticipated this.

This commit also corrects earlier comments of mine.  Several places described
these failures as "a stalled bus access rather than an abort", with the absence
of a recorded fault as the evidence.  That was wrong: they were ordinary
exceptions whose handlers were unreachable.  The two underlying bugs -- the
peripheral port disabled at reset, and the out-of-bounds MPU table write -- were
real and are correctly fixed, but the explanation for why they presented so
mutely was not.  Fixed in platform.h, mpu.c and bsp_boot.c.

Interrupt plumbing: gicv3.c is ported from the FVP example with the frame bases
in platform.h, derived from IMP_CBAR and TRM Table 9-1.  el1_irq_entry in
entry.S uses the classic A32 form rather than ThreadX's context save/restore,
since this image does not link the kernel.  irq_dispatch.c counts ticks instead
of calling _tx_timer_interrupt, and keeps separate spurious and unexpected-INTID
counters so that "no interrupt arrived" and "an interrupt arrived and was
mishandled" cannot be confused.  timer_start_oneshot_irq arms with IMASK clear;
the polled timer_start_oneshot keeps IMASK set, and the two are separate
functions because either mistake is silent.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Ran ThreadX on S32Z280 silicon with threads, tick and preemption

The kernel runs on the board:

    === ThreadX Cortex-R52 :: S32Z280-594EVB ===
    console   = 0x00000000
    entering kernel

    === ThreadX on S32Z280: results ===
    ticks     = 0x00000064      100 ticks
    sleeper   = 0x00000014       20 wakeups
    spinner   = 0x00159B48    1415496 iterations
    preempt   = 0x00000014       20 preemptions
    PASS threads, tick and preemption all verified

The figures are internally consistent, which is what makes them evidence rather
than encouragement: 100 ticks elapsed for a 100-tick sleep, so the tick rate is
exactly as configured; 20 sleeper wakeups is precisely 100/5 for a 5-tick sleep;
and 20 preemptions means every wakeup took the core from a runnable
lower-priority thread rather than receiving it by cooperative handoff.

This validates the context-switch assembly merged in #579 on real Cortex-R52
silicon.  Until now it had only run against the Armv8-R AEM FVP, which reports
MIDR part 0xD0F -- an architecture envelope model, not an R52 implementation.

The demo judges four things separately because on first silicon they fail
independently: that time advanced (the GIC delivers the timer PPI and
_tx_timer_interrupt is reached), that tx_thread_sleep returned (tick-driven
scheduling, not merely the interrupt), that the lowest-priority thread ran (the
sleeper actually yielded), and that the sleeper resumed while the spinner was
runnable (preemption).  A demo that only counted whether both threads ran would
pass with a broken tick if they happened to yield to each other.

Integration pieces, all from the FVP example with the board's differences:

  - tx_initialize_low_level.S publishes the system stack and the first free
    address, then calls board_init().  Shorter than the A-profile reference
    ports because entry.S has already given every mode its own stack from
    dedicated linker regions.

  - entry.S routes the IRQ vector through _tx_thread_context_save and
    _tx_thread_context_restore under TX_R52_USE_THREADX_IRQ, keeping the
    standalone A32 handler for s32z280_boot.elf, which does not link the kernel
    so that a boot failure there stays unambiguous.

  - board_init() runs mpu_init() FIRST.  The GIC distributor is unreachable
    until its region is mapped Device nGnRnE, so any GIC access before the MPU
    is enabled aborts.

  - link.lds provides _end as well as end; _tx_initialize_low_level publishes
    unused memory from _end.

  - bsp_done() is defined in the demo rather than shared from bsp_boot.c, whose
    bsp_main would collide with the demo's.

s32z280_boot.elf still builds and the FVP example still passes 6/6.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Tightened the S32Z280 MPU map and verified protection is enforced

The map now carries real permissions -- code read-only and executable, data
writable and never executable, peripherals Device nGnRnE and never executable,
everything else unmapped -- and the enforcement is demonstrated rather than
asserted:

    X1 write to read-only code region
    faults     = 0x00000001
    fault DFSR = 0x00000A0C      bit 11 (WnR) set: a write fault
    fault addr = 0x79900000      the address written
    X2 PASS write to code faulted and was recovered

Exactly one fault, identified as a write, at the precise address, followed by
successful resumption.  ThreadX still passes on the same map: ticks 100,
sleeper 20, spinner ~1e6, preempt 20, and IRQ count keeps advancing during the
protection test.

Reading SCTLR back or listing the programmed regions would only show what was
configured.  Provoking the violation shows it is enforced, which is the claim
that matters for a protection story.

entry.S gains a recoverable data-abort path, gated on a fault_expected flag the
test arms.  It records DFSR and DFAR, counts the fault, and resumes at the
instruction after the faulting access -- lr on data-abort entry is the faulting
address plus 8, so subs pc, lr, #4 skips the access instead of retrying it
forever.  Unarmed, a data abort stays fatal and reported, which is what a real
bug should get.

This supersedes the permissive bring-up map, and the reason that map existed is
worth recording: a narrow map had failed earlier and the cause was unknown, so
coverage was blamed.  Coverage was never the problem.  An out-of-bounds write to
a [3]-sized region table put a region at base 0 overlapping everything, and
SCTLR.TE being set meant the resulting abort could not reach a handler.  With
both fixed a narrow map works first time, and -- more usefully -- a mistake in it
now reports vector, syndrome and faulting address instead of hanging mutely.

Until TE was cleared no fault handler on this board could run, so no protection
claim about it could be tested at all.  This is the first one that could.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Enabled the S32Z280 caches, with an honest effectiveness result

Both caches are on and verified from SCTLR:

    CLIDR      = 0x09200003
    SCTLR      = 0x30C50839 -> 0x30C5183D    (C and I set)
    cachesOn   = 1
    counts off = 0x00094DF4
    counts on  = 0x00094CC8
    gain/1000  = 0
    C3 PASS caches enabled (SCTLR.C and SCTLR.I set)
    C4 no significant speedup -- expected here, the workload is already
       in fast local SRAM

cache.c supplies what the FVP example explicitly deferred to silicon: a
data-cache set/way sweep that reads CLIDR and CCSIDR and walks every set and
way the hardware reports, rather than assuming a geometry.  The FVP file notes
that this "belongs with the silicon bring-up where it can be verified against
the real cache geometry"; this is that.  The sweep matters because caches are
not architecturally guaranteed invalid out of reset, and enabling a write-back
data cache holding stale valid lines would evict them over live memory.

The effectiveness result is reported separately from the enable, and that
separation was earned.  The first version of this test passed on warm < cold and
printed "caches on and the workload got faster" for 609904 counts against
609686 -- a 0.036% difference that is measurement noise.  That criterion was
worthless and the PASS was misleading.  It now reports the gain as a fraction
and only claims a speedup above 10%.

No speedup here is a plausible result rather than a defect: both the code and
the data this workload touches live in RTU-local low-latency SRAM, close to core
speed already.  There is no slow memory on this board to demonstrate a cache
against -- DDR would be the place, and DDR is not initialised.  So the honest
claim is "enabled and harmless", not "enabled and beneficial".

cache_clean_all() is called before parking, and this is not optional.  With a
write-back data cache, values this image writes can sit dirty in cache where a
debugger reading SRAM cannot see them -- and the memory-resident progress
markers and fault records this bring-up depends on are read exactly that way.
Without the clean, a post-mortem read could report stale values and look like a
fault that never happened.

ThreadX still passes on the same configuration and the protection test still
faults correctly: write to the read-only code region takes exactly one write
fault at the written address and recovers.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Made the cache benchmark actually measure the cache

The caches now show a real effect instead of noise:

    D line     = 64 bytes
    D ways     = 4
    D sets     = 64
    D bytes    = 0x4000    16 KB L1 data cache, read from CCSIDR
    bench wds  = 0x800     8 KB working set, half the cache
    counts off = 0x000D1B87   858503
    counts on  = 0x0009F515   652053
    gain/1000  = 0xF0         24% faster
    C4 PASS caches measurably faster (>=10%)

The previous benchmark could not have measured anything, and its 0.036%
"speedup" was noise.  Two things were wrong with it.

It used a static buffer in the RTU-local fast-data bank.  That memory is close
to core speed already, so caching it saves almost nothing -- there is no latency
to hide.  The benchmark now runs over the extended SRAM at S32Z_EXT_SRAM_BASE,
which the Reference Manual describes as NOT RTU-local, and mpu.c gains a sixth
region to map it Normal write-back and never executable.

And its working set was a guessed 4 KB constant.  It is now sized at run time
from CCSIDR to half the reported data cache, so it fits and is revisited every
pass.  A set larger than the cache would stream through and evict, showing
little benefit even over slow memory -- a guess could have produced a null
result for a reason that has nothing to do with whether the cache works.

24% on this loop is a believable figure rather than a suspiciously large one:
the workload is a store-heavy read-modify-write over a write-back cache, so part
of the cost is write traffic that still has to reach SRAM.

Useful by-product: the L1 data cache geometry is now read and printed rather
than assumed.  Nothing in the SoC reference manual gives it, and the set/way
invalidate sweep in cache.c depends on it being right.

ThreadX still passes and the protection test still faults correctly.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Fixed the PRBAR field encoding and proved execute-never enforcement

program_region() shifted every PRBAR field one bit too far left: SH<<4,
AP<<2 and XN<<1, where the Cortex-R52 TRM (r1p3 figure 3-39, table 3-80)
places BASE[31:6], RES0[5], SH[4:3], AP[2:1] and XN[0].  The intended XN
therefore landed in AP[1] -- the EL0-access bit -- and no region was ever
execute-never.  On this board the core ran instructions straight out of
.data with LR_abt at 0x3178000c to prove it.

The write-to-read-only test passed throughout, which is why this survived:
under the old shift the AP value's low bit happened to land in the real
AP[2], the read-only bit, so permissions came out right by accident while
XN was silently discarded.  Only an instruction fetch from a region marked
non-executable could distinguish the two.

The AP macros in mpu.h go back to the architectural encoding, AP[2] for
read-only and AP[1] for EL0 access.  No calibration is needed; the values
were correct as published all along.

Verified on S32Z280 silicon: the data region now reads back PRBAR
0x31780001 with XN set, a write to the read-only code region faults with
DFSR 0xA0C, execution from the data region takes a prefetch abort with
IFSR 0x20C at 0x31780000, both faults recover, and the ThreadX demo still
reports 100 ticks with 20 preemptions.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Corrected the same PRBAR encoding bug in the FVP MPU example

The FVP example carries the identical off-by-one shift that the S32Z280
bring-up exposed, so no region there was execute-never either: the data
region and the whole peripheral region were both freely executable, and
each also gained unintended EL0 access from the misplaced XN bit.

This also retires the "PRBAR.AP bit order is reversed" claim in mpu.h.
That note rested on a real measurement -- four regions, one per AP
encoding, privileged write attempted on each, writes faulting for 0b01 and
0b11 -- but the cause was the shift, not the bit order.  With AP written
into bits[3:2], its low bit lands in the real AP[2], the read-only bit, so
writes fault exactly when that bit is set.  The "high bit grants EL0
access" half of the conclusion was never tested; under the old shift that
bit landed in SH[0], programming a shareability the TRM calls
UNPREDICTABLE.  The architectural encoding needs no calibration.

demo_mpu.c gains the check that would have caught this: an instruction
fetch from the execute-never data region must take a prefetch abort.
Recovering from one cannot work the way the data-abort path does, since
skipping the faulting instruction is impossible when the instruction is
what could not be fetched, so entry.S gains a recoverable prefetch-abort
path that returns to a landing point recorded by mpu_try_execute.

Verified both ways.  All five FVP images pass, and the new check reports
prefetch aborts 0 -> 1.  Restoring the old shift for one run makes it fail
with 0 -> 0, so the check is not vacuous.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Removed the unassemblable .arch directive from the S32Z280 example

GNU as accepts ".arch armv8-r"; LLVM's integrated assembler accepts no spelling
of it -- not armv8-r, not armv8r, not armv8-r+crc -- and stops with "Unknown
Arch: armv8-r". There is nothing to substitute, so the directive goes. The
architecture comes from -mcpu=cortex-r52 on the command line, which every
toolchain file passes, so it only restated it.

entry.S in this example never had one. The two equivalents in the FVP example are
handled separately, in the change that brings those images under the LLVM check.

Verified: both S32Z280 assembly sources now assemble with Arm Toolchain for
Embedded 22.1.0, and the GNU build of s32z280_boot.elf and s32z280_demo.elf is
unchanged with no warnings.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Stopped read_identity.gdb reporting failure after a successful demo run

The script reads r52_identity, which only the boot image defines. The kernel demo
shares the same bsp_done breakpoint and the same boot_stage marker, so it is
convenient to point the same script at either image -- but against the demo the
lookup raised and gdb exited 1, after the demo had already printed
"PASS threads, tick and preemption all verified" over the console. A tool that
reports failure on success is worse than one that prints nothing.

The lookup is now guarded, and says which image it is looking at rather than
falling over. Everything after it is skipped when the structure is absent.

Verified against both images on S32Z280 silicon: the demo exits 0 and reports its
own PASS, and the boot image still prints the full identity report with MIDR
0x411FD133 and C3, C4, X2 and X4 all passing.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-12 10:25:35 -04:00
Frédéric Desbiens b0ad67bae9 Brought the Cortex-R52 examples under the LLVM check, and fixed two blockers (#604)
#600 added a line reporting which example builds the LLVM check passes over, and
cortex_r52 was on it: its examples are driven by CMake rather than by a
build_threadx.sh pair, so the linking stage never touched them. Covering them
turned up two reasons they could not have been built with anything but GNU.

.arch armv8-r has no portable spelling. GNU as accepts it, and LLVM's integrated
assembler rejects every variant -- armv8-r, armv8r, armv8-r+crc -- with "Unknown
Arch: armv8-r". There is nothing to substitute, so the directive is gone from
entry.S and tx_initialize_low_level.S; -mcpu=cortex-r52 already selects the
architecture, both toolchain files pass it, and the directive only restated it.
Worth noting where this hid: the assembly stage walks ports/*/gnu/src, so example
assembly had never been assembled by LLVM at all.

-Wl,--no-warn-rwx-segments is GNU ld only, added in binutils 2.39. ld.lld does
not warn about RWX segments and rejects the flag outright, failing the link with
"unknown argument". It is now selected on CMAKE_C_COMPILER_ID rather than spelled
into all six targets, and the reason it exists at all -- a bare-metal image has
one flat DRAM region and leaves access control to the MPU -- moves to the one
place that sets it.

cmake/cortex_r52_clang.cmake is the toolchain file. It names the tools as found
on PATH, which is what CI uses, then pins $HOME/toolchains if that directory
exists, mirroring how cortex_r52.cmake pins the GNU toolchain and for the same
reason. Falling back rather than requiring the pinned path keeps the file usable
on a machine that keeps clang elsewhere. THREADX_TOOLCHAIN stays "gnu": there is
no clang port directory, this builds the gnu sources with a different compiler,
which is what the whole check does for every other Arm port.

check_clang.sh gains a fourth stage for CMake-driven examples, and no longer
reports cortex_r52 as a gap. It reads the image list out of the generated ninja
graph rather than repeating it, so adding a target cannot escape the check, and
filters out the cmake_object_order_depends_target_* phonies -- counting those
reported ten images where there are five.

Verified with Arm Toolchain for Embedded 22.1.0, the version the workflow pins.
All five images link, and all five then run and pass on FVP_BaseR_AEMv8R:
boot_check, demo_m2, demo_m3, demo_threadx and demo_mpu. That is a step beyond
the AArch64 examples, which are link-verified only. The full check reports 711 of
711 assembly sources, 185 of 185 common C sources for each of nine cores, 42 of 42
script-driven examples and 5 of 5 CMake images, leaving only cortex_a5_smp,
cortex_a7_smp and cortex_a9_smp listed as having no driver.

GNU is unaffected: the same five images build with no warnings and the FVP test
suite passes 5 of 5.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-11 19:53:11 -04:00