Commit Graph
490 Commits
Author SHA1 Message Date
Frédéric Desbiens d55fe1ef2f Stopped the release script adding a Co-authored-by trailer (#777)
prepare_release.sh committed the version passes with a Co-authored-by trailer naming an
AI. That trailer asserts authorship an AI cannot hold: the human contributor signs the
ECA and is solely responsible for the contribution. It is already in the published
history, on the 6.5.1.202602 and 6.5.1.202602a preparations.

The commits now carry their subject alone. A version pass is mechanical sed output, so
no agent produces it at run time; an agent that runs the script records its own
Assisted-by trailer on that run instead. The script also gains the AI disclosure line it
was missing.

Ran the patched script against a scratch clone targeting 6.5.2.202603. Both commits come
out with an empty trailer block, the version constants and 211 port version strings
update as before, and check_ports.sh passes.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-28 07:28:23 -04:00
Frédéric Desbiens b37cd4a81a Fixed RISC-V regression portability failures (#773)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
RISC-V regression builds ran only at -O0, leaving a timer callback counter that
spins forever at -O2. The trace regression was excluded because it referenced
a port-specific interrupt-save variable, and its adjacent pools could start
misaligned on RV64.

Made the timer counter volatile, gave the trace test aligned pool storage and a
portable saved-interrupt value, and enabled it on RISC-V. Added an -O2 QEMU
configuration to keep the optimized failure covered.

CMake/Ninja/QEMU: RV64 default, optimized and trace suites passed 97/97 each;
RV32 passed 96/96 each. Two ISR event tests passed 30 repeats each at -O2.
The Linux/GCC 14 trace test passed. The reported timing resonance did not recur.

Assisted-by: Codex (gpt-6-sol) <noreply@openai.com>
2026-09-23 16:52:44 -04:00
Frédéric Desbiens 70a5300977 Added a ThreadX module manager port for the Cortex-R52 (#639)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
r52_fvp / r52 (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
Added the first GNU ThreadX module port for an Arm R-profile core.

  The port combines the PMSAv8-R MPU model from Cortex-M33 with the AArch32
  privilege and processor-mode handling from Cortex-R4. It provides the complete
  module path: headers, manager sources, scheduler integration, user-mode entry,
  SVC dispatch, data- and prefetch-abort capture, fault notification, relocatable
  module loading, and shared-memory regions.

  Module isolation uses eight MPU regions (8–15) with a 64-byte granule. The
  scheduler replaces those regions on each module switch and manages a separate
  privileged loading window in region 16. Assembly-visible structure offsets and
  region-layout assumptions are checked at build time.

  Added independent module demonstrations for the S32Z280-594EVB and Armv8-R AEM
  FVP. The automated FVP regressions cover:

  - Loading the same position-independent module at different addresses
  - User-mode data and instruction access violations
  - Fault capture and notification for both abort types
  - Shared-region access, alignment, exhaustion, empty-size, and overflow handling
  - Required module-property combinations
  - The GCC CLZ-based priority search

  The port requires user mode and memory protection together, at least 17 EL1 MPU
  regions, and currently validates A32 modules; Thumb module execution remains
  unvalidated.

  Validated with GNU Arm 14.3.1. The FVP suites pass 8/8 in the default
  configuration and 11/11 with hard-float, FIQ, and interrupt nesting enabled.
  The S32Z280 images build cleanly, and the module isolation and relocation paths
  were exercised on S32Z280 silicon during development.

  Matching user documentation is provided by rtos-docs-asciidoc PR #41.

  Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-17 15:39:30 -04:00
Frédéric Desbiens ad558a7b1f Fixed the zero trace time stamps in the Linux ports' MISRA builds (#749)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
Both Linux ports define TX_TRACE_TIME_SOURCE as _tx_misra_time_stamp_get() when
TX_MISRA_ENABLE is set, and neither implements that function, so both inherit
the generic `return(0);` from tx_misra.c. Every trace event is stamped zero.
The buffer carries no timing, and the kernel's own check for an entry having
been overwritten -- time_stamp against the entry's own stamp, in the block and
byte allocates and in the system suspend and resume -- compares zero with zero,
so it never fires and a service patches whatever now occupies the slot.

Both ports now read in MISRA builds the clock they already read otherwise,
_tx_linux_time_stamp.tv_nsec, which TX_TRACE_PORT_EXTENSION refreshes on every
recorded event in both forms of the insert. The non-SMP port's non-MISRA macro
carried a trailing semicolon, which made it a statement and is why the MISRA
insert -- which takes the time source as a function argument -- could not use
it; that is dropped and the two branches become one definition. Both headers
keep the _tx_misra_time_stamp_get declaration, because tx_misra.c still defines
it and is compiled for these ports.

The MISRA insert evaluates its time source before the callee refreshes the
clock, so each entry carries the reading taken at the previous recorded event.
Stamps are real, distinct and ordered, which is what the overwrite check needs.

The trace entry update test gains an assertion that the buffer holds an entry
the port actually stamped. It fails on dev with ERROR #13 under
misra_trace_build and passes with this change. Suites green: 7/7 ThreadX
configurations, 5/5 SMP.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-16 14:01:07 -04:00
Frédéric Desbiens e9b5aec65e Added MISRA build configurations to the ThreadX regression matrix (#747)
TX_MISRA_ENABLE routes the kernel's pointer conversions, its memset and its
stack check through the shim in common/src/tx_misra.c. No build configuration
defined it, so that file was never compiled or tested, and the source the MISRA
analysis runs against was not the source the suite exercises. Adding event
tracing alongside it reaches a further region of the shim that neither macro
alone compiles. Both defects found in this area were invisible for the same
reason: #741 broke under TX_MISRA_ENABLE and #745 under both macros together,
and neither combination was built anywhere.

misra_build and misra_trace_build fill that gap. Two things had to give way for
them. TX_POINTER_TO_ALIGN_TYPE_CONVERT exists only when TX_MISRA_ENABLE is
absent, so threadx_test_port.h writes the conversion out rather than taking it
from the API; it is test scaffolding storing a pointer in a word wide enough to
hold one, not kernel source. And thread_transition and module_manager compile
hand-picked common/src sources directly instead of linking the library, which is
what lets them reach feature macros the library-per-configuration model cannot;
under TX_MISRA_ENABLE those sources call into the shim, and pulling the shim in
pulls its own callees after it. Neither directory exists to test the shim, so the
MISRA configurations skip them.

100/100 tests pass in each of the two new configurations, against 105/105 in the
existing five - the five not run are the four thread transition tests and the
module manager test, exactly the two directories skipped. Build is 4 to 5
seconds and the suites 13 and 15 seconds, against 4 seconds and 9 to 17 for the
configurations already there, so the tx job grows by roughly forty seconds.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-16 09:55:24 -04:00
Frédéric Desbiens 24d1f278c2 Decorrelated the wait abort ISR test's interrupt source from the thread it samples (#743)
threadx_thread_wait_abort_and_isr_test waits for the periodic timer interrupt to
land while the preempt disable flag is set. The same interrupt also puts the
semaphore that wakes the thread being sampled, so the two run in lockstep: the tick
wakes thread 0, thread 0 does a fixed amount of work and suspends, and the next tick
arrives a fixed interval later with thread 0 in the same place every time. That is
the resonance the handler's own comment describes, and perturbing the handler's
duration only shifts the phase rather than breaking the correlation. The window
itself is a few instructions wide, so a sampler locked to the thread's own cycle can
miss it indefinitely, which is what #644 and #649 measured and worked around.

The simulator's timer thread waits on _tx_linux_timer_semaphore with a one-tick
deadline and delivers an interrupt early when the semaphore is posted, which the port
already relies on in _tx_thread_schedule. The test now runs a plain POSIX thread that
posts it, injecting interrupts at moments unrelated to the tick grid and sampling
thread 0 at arbitrary points in its cycle rather than the same one. The injector posts
only when nothing is outstanding, so interrupts can never be queued faster than they
are serviced. It is confined to the Linux simulation port and no port file changes.

The count of windows asked for goes back to ten, on the same reasoning that lowered
it to three: ask for what a run can actually reach. Measured over six build
configurations of a branch carrying this change, every one reaches ten of ten in under
a second, against one to three of three previously with three of the six pinned at the
180 second budget. The budget and the zero window ceiling are untouched, so a run that
somehow still falls behind behaves exactly as it does today.

The SMP copy deliberately does not get the injector, and its count stays at twenty. It
has never been in the slow mode, and injecting interrupts there measurably hurts:
8.9 seconds against 0 to 1 for the same twenty windows, which is the extra interrupt
traffic contending across the simulated cores. Only the tx copy has the problem this
solves, so only the tx copy changes; the budget and ceiling logic stays identical
between them.

Verified on this branch: the tx suite passes 105 of 105 with the test at 0.44 seconds,
and four direct runs reach ten of ten in 0 to 1 seconds. The SMP suite passes 117 of
117 unmodified, with its own test at 0 to 1 seconds over four runs.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-16 09:18:41 -04:00
Frédéric Desbiens b93d1ee92f Fixed the simulator ports and thread create paths so they compile when TX_MISRA_ENABLE is defined (#742)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
r52_fvp / r52 (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
* Fixed the simulator ports so they compile when TX_MISRA_ENABLE is defined

_tx_thread_stack_build() in the four simulator ports converts the fake stack
pointer through TX_POINTER_TO_ALIGN_TYPE_CONVERT and
TX_ALIGN_TYPE_TO_POINTER_CONVERT. tx_api.h defines both macros only in the
non-MISRA branch of its #ifdef TX_MISRA_ENABLE, so with that macro defined the two
names are undeclared and none of the four files compiles.

Reproduced with:

    gcc -m32 -c -DTX_MISRA_ENABLE -I common/inc -I ports/linux/gnu/inc \
        ports/linux/gnu/src/tx_thread_stack_build.c -o /dev/null

which reports both names as implicit declarations and then an int to pointer
assignment. The same command without the define compiles cleanly.

No build configuration under test/tx/cmake or test/smp/cmake defines
TX_MISRA_ENABLE, so CI never compiles these files in that mode. It surfaced on a
branch that carries such a configuration.

The conversions are now written inline, which is what the non-MISRA macros expand
to and what the surrounding port code already does, including the line this
replaced.

The alternative would be the idiom common/src/tx_thread_create.c uses for the same
conversion: an explicit #ifdef selecting the ULONG pair under MISRA. That is not
equivalent here. On __x86_64__ this port defines ULONG as unsigned int and
ALIGN_TYPE as unsigned long long, so a pointer round-tripped through the ULONG pair
loses its top 32 bits. Writing the conversion inline keeps one form that is correct
in both modes and on both widths.

Verified by compiling the Linux port with and without TX_MISRA_ENABLE, both clean.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Fixed the same MISRA build break in the SMP and module manager thread create

_tx_thread_create() in common_smp and _txm_module_manager_thread_create() convert the
thread's stack start through TX_POINTER_TO_ALIGN_TYPE_CONVERT and
TX_ALIGN_TYPE_TO_POINTER_CONVERT with no conditional at all. tx_api.h defines both
macros only in the non-MISRA branch, so with TX_MISRA_ENABLE and
TX_ENABLE_STACK_CHECKING both defined neither file compiles. It is the same defect as
the simulator ports in the previous commit, in two more files.

Reproduced with:

    gcc -c -DTX_MISRA_ENABLE -DTX_ENABLE_STACK_CHECKING \
        -I common_smp/inc -I ports_smp/linux/gnu/inc \
        common_smp/src/tx_thread_create.c -o /dev/null

which reports both names as implicit declarations. Both files now compile with and
without TX_MISRA_ENABLE.

The conversions are written inline for the same reason as the ports: the #ifdef idiom
that common/src/tx_thread_create.c uses selects the ULONG pair under MISRA, which
truncates a pointer wherever ALIGN_TYPE is wider than ULONG. That is reported
separately.

Verified: the SMP regression suite passes 117 of 117, and the ThreadX suite still
builds.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-15 23:19:32 -04:00
Frédéric Desbiens 26e4aa182c Fixed the build of tx_misra.c with MISRA and event trace enabled (#746)
tx_misra.c defines five pointer conversions for the trace component, and their
prototypes sit in tx_trace.h behind TX_SOURCE_CODE. tx_misra.c is the one kernel
source that does not define that macro, so it defines all five with no previous
declaration. Under -Wmissing-declarations -Werror, which the kernel target is
built with, the file does not compile at all, so TX_MISRA_ENABLE and
TX_ENABLE_EVENT_TRACE cannot be enabled together. No build configuration defines
both, which is why this has never surfaced.

The five prototypes move out of the TX_SOURCE_CODE guard into their own block,
next to where tx_api.h already declares _tx_misra_trace_event_insert - the sixth
function of the same region, which escapes the problem for exactly that reason.
Declarations only; no definition moves and no macro changes meaning.

All 185 files in common/src now compile clean under all four combinations of the
two macros with -Wall -Wextra -Wmissing-declarations -Werror, against five errors
in tx_misra.c for the both-enabled case before. Object code is byte identical for
the three combinations that already built: 0 of 185 files differ in each.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-15 23:08:23 -04:00
Frédéric Desbiens 9e4c57138d Normalized the AI disclosure comment to one fixed line per file (#740)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
r52_fvp / r52 (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The per-edit disclosure named the product and model, so every agent and every
model version appended another line. 74 files carried two to four of them, and
the same five products had accumulated 13 spellings -- Copilot against GitHub
Copilot, Claude Sonnet 4.6 against claude-sonnet-4.6, four spellings of Codex.
Twenty assembly lines carried a doubled comment marker, `; //` or `@ //`.

Every file now carries exactly one line, fixed text naming no product:

    Portions of this file were generated with AI assistance.

It is written with the comment character that file already uses, so the `;`
and `@` assembly files keep theirs and the doubled markers are gone. Precise
attribution stays on the commit, where the Assisted-by trailer is per-change,
dated and attached to the diff it describes. A header line cannot hold that
record honestly, because the code it names gets rewritten and the line stays.
A file-level flag answers whether; the history answers who.

Comment-only. 455 files, 455 insertions and 574 deletions: every removed line
was a disclosure line, every added line is the fixed text, and no file is left
with zero or with more than one. `scripts/check_ports.sh` passes, including the
reproducibility check that would catch a ports_arch master and its generated
copies drifting apart. Recompiled against dev, every file that builds without a
vendor toolchain gives a byte-identical object: 19 of 19 C files under common,
100 of 100 GNU assembly files, and all 16 assemblable files whose comment
marker changed. The 10 remaining marker changes are ac5 and IAR sources where
`;` already started the comment and only the redundant `//` was removed.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-15 17:17:26 -04:00
Frédéric Desbiens f5e59a1dd3 Completed Windows simulator support and regression coverage (#736)
Completed and stabilized the Win32 and Win64 MSVC simulator ports.

- Replaced high-latency host synchronization with bounded scheduler handoffs and
  critical sections.
- Added a high-resolution timer, tick batching, idle fast-forward, shutdown
  coordination and Win64 extension-pointer support.
- Brought the Windows regression tooling and the new thread-transition tests up
  on CMake, Ninja and the Visual Studio Build Tools.
- Extended the SMP teardown diagnostics and corrected 64-bit trace-test handling.

This supersedes the historical `win64`, `win32-perf`, `windows-sim-ports` and
`windows-sim-ports-completion` branches; no unmerged change from them is missing
here. The original Win64 port landed in #529.

1,610 of 1,610 tests pass, across five configurations each: Win32 515, Win64 515,
Win64 SMP 580. No external dependency was added, and the existing MSVC warnings
in the trace configuration are unchanged. Hardware validation does not apply to
host simulator ports.

Assisted-by: Codex (gpt-5.6-sol) <codex@openai.com>
2026-09-15 16:47:08 -04:00
Frédéric Desbiens 30b3d22adc Restored the random stack fill value cleared during thread creation (#732)
Fixes #723

With `TX_ENABLE_STACK_CHECKING` and `TX_ENABLE_RANDOM_NUMBER_STACK_FILLING` both
enabled, `_tx_thread_create` picked a random byte, stored it in
`tx_thread_stack_fill_value`, filled the stack with it -- and then cleared the
whole control block with `TX_MEMSET`. The stack held the pattern while the
control block claimed zero, so `TX_THREAD_STACK_CHECK` and
`_tx_thread_stack_analyze` compared against the wrong value for the entire life
of the thread.

The value is now computed into a local and written back after the clear. The
clear stays where it is, because the module manager's error checking walks the
created list before the control block may be touched.

Fixed in `common/src/tx_thread_create.c`, `common_smp/src/tx_thread_create.c`
and `txm_module_manager_thread_create.c`, where it additionally left a user-mode
module thread's kernel stack filled with zeros.

`threadx_thread_stack_fill_value_test` creates sixteen unstarted threads and
checks each control block against the pattern in its stack, tolerating a random
zero byte without letting that hide the defect. Added to both suites: `ERROR #3`
on `dev` in `stack_checking_rand_fill_build`, green with the fix. Five tx
configurations at 104 tests, five SMP at 117, and all three sources clean under
`-Wall -Wextra` across every combination of the four stack-filling switches.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-15 16:24:58 -04:00
Frédéric Desbiens 201cc04609 Applied R_ARM_RELATIVE relocations at GNU ThreadX module startup (#731)
Fixes #230

`gcc_setup.s` rebases the GOT and copies initialized data into the module data
area, but never applies the relocations the linker records for words inside that
data. A global initialized with a link-time constant -- `char *pBuffer = buffer;`,
a function pointer, a struct member holding a string literal's address -- kept
its link address and pointed outside the loaded module.

`gcc_setup.s` now walks `.rel.dyn` after the data copy and adjusts every
`R_ARM_RELATIVE` entry targeting the module data area, reusing the GOT loop's
base arithmetic and skip rules, and reloading the flash base first because
`crt0_memory_copy` clobbers `r3`.

The table is only emitted with `-pie`, so the example build scripts now pass
`-pie --no-dynamic-linker` and the example linker scripts discard `.dynamic`,
which `-pie` would otherwise place at address zero and inflate the image to
nearly 200 KB. Without `-pie` the table is empty and the loop is a no-op.

Covers Cortex-M3, M4, M7, M33 and M0+. Cortex-M23 ships no module linker or
build script and AArch64 uses a different relocation format; both left for a
follow-up.

Verified on `qemu-system-arm` with a module linked at one address and loaded at
another: char, function, string and struct-member pointers all resolve to the
loaded addresses, and the same module built without `-pie` still runs.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-15 16:24:54 -04:00
Frédéric Desbiens 3e5a16c95b Allowed tx_timer_change to be called from tx_application_define (#730)
Fixes #224

`_txe_timer_change` returned `TX_CALLER_ERROR` for any call at or above
`TX_INITIALIZE_IN_PROGRESS`, making `tx_timer_change` the only timer service
that could not be called from `tx_application_define` -- while still being
allowed from an ISR.

The check has no technical basis: `_tx_timer_change` only writes the expiration
fields of a timer that is not on an active list, with interrupts disabled.
Removed it from both the `common` and `common_smp` copies of
`txe_timer_change.c`, with the now-unused includes and the `TX_CALLER_ERROR`
line in the header comment. Relaxing an error check is backward compatible.

`testcontrol.c` in both suites now calls `tx_timer_change` at initialization, so
the timer simple test covers it: `ERROR #30` on `dev`, green with the fix.
116/116 SMP, 103/103 non-SMP.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-15 16:24:50 -04:00
Frédéric Desbiens 9b2979e6b0 Replaced the GNU-only dsb/isb 0xF operands with the UAL sy form (#729)
Fixes #551

`_tx_thread_system_return_inline()` in the Cortex-M `tx_port.h` headers spells
its barriers `dsb 0xF` and `isb 0xF`. A bare hexadecimal operand is a GNU
assembler extension, and IAR rejects it with `operand syntax error`, so the
header cannot be included at all. The block is guarded for GCC, armclang and IAR
together, so every IAR user of an affected port hits it -- four independent
reports on Cortex-M33 and M7 with EWARM 9.50 and 9.70.

Both operands become `sy`, the Arm UAL name for exactly what `0xF` encodes. The
generated instruction is unchanged. Applied to the two `ports_arch` masters and
all 32 copies under `ports`, covering M0, M23, M3, M33, M4, M52, M55, M7 and M85
across ac5, ac6, gnu, iar and keil, plus the `scripts/check_ports.sh` probes that
matched the old spelling.

`check_ports.sh` passes, the copy scripts still reproduce every generated port
byte for byte, and `arm-none-eabi-gcc -O2` compiles a caller for every patched
header, emitting `dsb sy` and `isb sy`. Three headers that need toolchain
intrinsics GCC does not ship fail identically on `dev`.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-15 16:24:45 -04:00
Frédéric Desbiens d15f28ab9f Masked the SMP remap core maps to silence a false -O2 array bounds error (#728)
Fixes #469

`_tx_thread_smp_remap_solution_find` indexes `_tx_thread_smp_schedule_list` with
the lowest set bit of the core maps it is given. Every caller already masks
those maps with `TX_THREAD_SMP_CORE_MASK`, but the compiler cannot see it, so
when the function is inlined into `_tx_thread_system_suspend` at `-O2` GCC
assumes a bit number as high as 31 and reports an out-of-bounds subscript.
`-Werror` turns that into a build failure.

Masked the three incoming maps at the top of the function, in both the inline
version in `common_smp/inc/tx_thread.h` and the twin in
`tx_thread_smp_utilities.c`. The masks are semantic no-ops, so scheduling is
unchanged, but the range is now visible to the optimizer. The first core queue
entry is also initialized, because the narrowed range lets GCC consider an empty
queue and warn about that instead.

Reproduced on the Cortex-A9 SMP port with arm-none-eabi-gcc 13.2.1, and clean
afterwards across -O2, -O3 and -Os and 2, 4 and 8 core configurations. SMP suite
116/116.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-15 16:24:40 -04:00
Frédéric Desbiens 1d4a4aeec8 Guarded _tx_thread_stack_analyze against inverted stack pointers (#727)
Fixes #460

`TX_ULONG_POINTER_DIF` casts the pointer difference to `ULONG`, so when
`tx_thread_stack_highest_ptr` sits below `tx_thread_stack_start` the midpoint
wraps to a huge value, the probe lands outside the stack, and the search never
converges: the caller hangs or faults.

`_tx_thread_stack_analyze` now requires the highest pointer to be strictly above
the start of the stack, and bounds the final scan by it. #464 covered the
`TX_THREAD_STACK_CHECK` path; this covers direct callers too, as
@billlamiework suggested on the issue.

Inconsistent pointers still mean a real overflow or a corrupted control block,
which remains the application's problem. What changes is that ThreadX reports it
through the stack error handler instead of hanging.

New cases for an inverted and an equal pointer pair crash the suite without the
fix and pass with it.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-15 16:24:35 -04:00
Frédéric Desbiens 5d235a534c Fixed the garbage _tx_initialize_unused_memory in the GNU Cortex-A ports (#726)
Fixes #435

`LDR x1, =__top_of_ram` already loads the top of RAM, so the `LDR x1, [x1]`
that followed read whatever sat at that address and left
`_tx_initialize_unused_memory` holding garbage.

Dropped that instruction from the 13 non-SMP GNU Cortex-A ports, the ARMv8-A
source they are generated from, and the two Cortex-A35 module examples. The SMP
GNU ports already had the correct form, and the Arm Compiler ports are
unaffected -- their symbol really does need the dereference.

`scripts/check_ports.sh` passes and the `gnu` CI job build-verifies the AArch64
ports. Not run on hardware.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-15 16:24:09 -04:00
Frédéric Desbiens cd6a2d9034 Put the module manager test's thread on the kernel's created list, so the stand-in kernel matches the one the manager reads (#739)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The host test built a TX_THREAD, gave it an ID and a module instance, and
left it off the kernel's created list. A created thread lives on that list,
so the thread the test created was not one the kernel would recognise.

The created list head and count are now defined for each object type, the
thread the test creates goes onto the thread list the way thread create puts
it there, and a successful delete takes it off again the way thread delete
does. Only the thread list is populated; the others stay empty because
nothing here creates an object of those types.

No expectation changed.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-14 17:09:00 -04:00
Frédéric Desbiens 83a3e62c28 Added a Module Manager host test harness and released a module thread's kernel stack only when it has one (#738)
* Released a module thread's kernel stack only when it has one, so deleting a thread without a manager-allocated stack no longer asks for that memory back

The thread delete dispatcher decided whether to release a kernel stack from
the calling module's property flags. A user-mode module can be given the
address of a thread that carries no kernel stack the manager allocated, and
deleting it passed that thread's null stack pointer to
_txm_module_manager_object_deallocate().

The pointer is now tested instead of the module's properties. Both thread
create paths clear the whole control block before filling it in, so a thread
with no manager-allocated kernel stack holds TX_NULL in the field, which
makes the test exact rather than an inference from how the module was
loaded.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Added a Module Manager host test harness and covered the user-mode thread kernel stack lifetime

The Module Manager is not in the ThreadX library the test tree links, and its
control blocks come from a module port rather than a base port, so nothing in
the suite could reach it. The thread_transition directory already describes
this technique as the one "the module manager tests in this tree use", but no
such tests existed.

This adds the directory that comment refers to: a port shim that replaces the
three interrupt primitives with host equivalents that count, and a CMake
target that compiles the manager sources under test directly against one
module port's headers.

The dispatch layer is one header of static functions and an unoptimised build
emits all of them, so a test that includes it to reach one dispatcher would
pull in references to every service the manager can dispatch. The header
guards each dispatcher with its own TXM_<SERVICE>_CALL_NOT_USED macro, so the
build reads that list out of the header and defines every guard except the
ones the test needs. Reading it rather than writing it down means a service
added later is excluded without anyone having to remember.

The first test asserts the invariant a user-mode module thread has to hold:
one logical thread costs the object pool two allocations, the control block
the module asks for and the kernel stack the manager takes on its behalf, and
deleting the thread must return the pool's available bytes and the module's
allocation-list count to exactly what they were. It asserts an equality
rather than a bound, because a bound would pass while one of the two was left
behind on every cycle, and then runs enough cycles that a leak of one stack
per cycle exhausts the pool several times over.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-14 16:08:48 -04:00
Frédéric Desbiens dde43b8ab2 Replaced the stale system stack switch pseudo-code in the ARMv7-A ports with comments that describe what the code actually does (#735)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The context save, vectored context save and system return routines in the
ARMv7-A ports carried pseudo-code comments claiming that they saved the
thread stack pointer and then switched to _tx_thread_system_stack_ptr.
Neither of those things happens, and none of these ports references that
variable outside an unused IMPORT in their example builds.

On ARMv7-A each processor mode has its own banked stack pointer. The IRQ
handler branches to _tx_thread_context_save while still in IRQ mode, so the
core's banked IRQ stack already serves as the system stack, and the thread
stack pointer is stored in the control block by _tx_thread_context_restore,
and only when the interrupt results in preemption. The scheduler runs on the
banked SVC mode stack that the startup code sets up. There is nothing for a
software stack switch to do.

The comments were therefore misleading rather than merely redundant, and had
led at least one user to try to restore the code they described. They are now
replaced by a description of the actual mechanism.

The AArch64 SMP ports keep their comments unchanged, because ARMv8-A does not
bank a stack pointer per processor mode and those ports do reload
_tx_thread_system_stack_ptr[core] explicitly.

This is a comment-only change. Every changed line is a comment, and all
twenty-five GNU variants still assemble cleanly for their target core.

The fourteen files under ports/cortex_a{5,7,8,9,12,15,17} were regenerated
from ports_arch/ARMv7-A/threadx/common/src/tx_thread_system_return.S with
ports_arch/ARMv7-A/update.sh. The ARMv7-A SMP ports have no generator, so
those files were edited directly.

Fixes #734

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-11 13:27:32 -04:00
Frédéric Desbiens c0aa4dbe29 Initialized the suspend status before the non-interruptable suspend in tx_thread_sleep, so a sleep no longer returns a stale error left over from an earlier timed-out suspension (#725)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
r52_fvp / r52 (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
When TX_NOT_INTERRUPTABLE is defined, _tx_thread_sleep did not set
tx_thread_suspend_status to TX_SUCCESS before calling
_tx_thread_system_ni_suspend. The interruptable path always did.

Nothing else writes that field on behalf of a sleeping thread:
_tx_thread_timeout resumes a TX_SLEEP thread directly and there is no
suspend cleanup routine for sleep. The value therefore survived from
whatever suspension the thread performed last, and tx_thread_sleep
returned it. A tx_semaphore_get that timed out with TX_NO_INSTANCE
would be followed by a tx_thread_sleep that also returned
TX_NO_INSTANCE after sleeping correctly.

The same change is applied to the SMP copy, which is identical.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-10 12:06:54 -04:00
Frédéric Desbiens 70741198e9 Refused thread delete and reset while an exit transition is in progress (#724)
_tx_thread_shell_entry and _tx_thread_terminate both publish a thread's
terminal state -- TX_COMPLETED or TX_TERMINATED -- and then call that
thread's exit notification callback, before the thread has been detached
from the ready list and before either service has finished with the pointer
it holds to the control block. That terminal state is exactly the state
_tx_thread_delete and _tx_thread_reset accept as authorization to
invalidate or rebuild the control block, and neither service tested whether
the transition producing it had finished.

A callback could therefore delete the thread it was called for -- and then
lawfully recreate it over the same memory, since delete exists to permit
that -- while the scheduler was still linked to the old incarnation.
tx_thread_create zeroes the whole control block and can auto-start the new
one, so the old priority list is left heading at a block whose own priority
field names a different list, with the old priority's map bit set behind
nothing. Alternatively a callback could reset a terminated thread, which
moves it out of the terminal state, and then resume it: the interrupted-
suspension logic in _tx_thread_system_resume refuses to void a suspension
only while the state is still terminal, so with the reset allowed first the
resume clears the suspending flag and restores TX_READY, the outer service
then finds the flag clear and skips the removal, and tx_thread_terminate
returns TX_SUCCESS for a thread that is runnable again.

Interrupt masking does not close the window, because the kernel restores
the prior posture before invoking the callback deliberately. On SMP
TX_RESTORE also releases the global protection, so the target can be
executing on another core while its callback runs -- and a reset there
memsets the stack a live core is running on. On the Linux, Win32 and Win64
host simulation ports the consequence is more immediate than corruption:
TX_THREAD_DELETE_PORT_COMPLETION cancels and joins the host thread backing
the deleted thread, so a callback-side delete of the completing thread
destroys the host thread the callback is running on.

The fix marks the transition and has the two services refuse a marked
target, returning the errors they already document, TX_DELETE_ERROR and
TX_NOT_DONE. The refusal is transient and the same call succeeds once the
transition has completed, so no documented lifecycle is lost; and it is in
the core services rather than the _txe_ wrappers, so disabling error
checking cannot disable it. Refusing the reset is also what closes the
resume path, without touching _tx_thread_system_resume: its existing
terminal-state test is sufficient once nothing can turn the terminal state
into TX_SUSPENDED from inside the window.

tx_thread_suspending is the marker, rather than a new control-block field.
It already means "a suspension is in progress" and is already true across
the callback in the two interruptable paths, so no field is added, the
public structure is unchanged, and sizeof(TX_THREAD) is unchanged --
which matters, because the Module Manager's object handling depends on the
sizes of the control blocks. Widening its lifetime was checked against
every reader rather than assumed. There are four: two in
_tx_thread_system_suspend and two in _tx_thread_system_resume. In every
window this change widens, the state is TX_COMPLETED or TX_TERMINATED, and
both resume readers already refuse to void a suspension for exactly those
two states, so their behaviour is unchanged; and no suspension routine is
called on the target in those windows, so the suspend readers never see
them. No suspension-initiating service can set the marker again inside a
window either: every one of them acts on a thread that is ready or
suspended.

Three sites needed changing beyond the two refusals, and the shape of each
was decided by where the marker can safely be cleared:

  - The non-ready branch of _tx_thread_terminate cleared the marker before
    the terminated extension and the callback, which is what left them free
    to act on a control block the service still had mutex-release
    processing to do against. The clear moves to the common tail, after the
    last dereference of the target, and becomes the single clear site for
    the whole service. In the interruptable ready branch the flag is
    already false there, because _tx_thread_system_suspend cleared it when
    it detached the thread, so the tail store is a second store of a value
    the flag already holds -- cheaper than testing for it, and it keeps one
    clear site.

  - Under TX_NOT_INTERRUPTABLE neither path set the marker at all, because
    that configuration does not use the interruptable suspension path that
    sets it. Both now set it before the callback. Interrupts being disabled
    there does not help: the callback is reached by a direct call.

  - In the TX_NOT_INTERRUPTABLE completion path the marker is cleared
    before _tx_thread_system_ni_suspend rather than after it. That call
    returns to the scheduler for a thread that is the current thread, which
    a completing thread is, and does not come back; clearing afterwards
    would leave a normally completed thread marked for ever and therefore
    permanently undeletable. Nothing is lost by clearing early there,
    because everything from that point to the detachment runs with
    interrupts disabled and calls no application code.

The change is the same change twice. All four files are byte-for-byte
identical between common and common_smp at this commit and stay so after
it, so common_smp was written by copying rather than by repeating the
edits. tx_thread_system_suspend.c and tx_thread_system_resume.c, which do
differ between the kernels, are deliberately untouched.

Tests. The in-tree regression test goes to both trees and is byte-for-byte
identical between them. It drives seven scenarios: terminating a ready
non-current target with two peers ready at the same priority, with the
callback attempting the delete and recreating the block if it succeeded;
the same with the callback attempting the reset and then the resume;
terminating a target suspended on a semaphore while owning a mutex, which
is the non-ready branch; natural completion alone at its priority,
including the safe post-completion reset, terminate, delete and recreate at
another priority; self termination; a benign callback, whose notification
count and ordering are unchanged; and the state and boundary cases, where
the new refusal must not fire.

It measures rather than describes. The callback records the published
state, the marker, and the status of every lifecycle service it can reach,
calling the core service as well as the wrapper wherever a refusal is
expected. A snapshot taken under interrupt lockout -- which is the global
SMP protection on an SMP port -- checks that every ready list agrees with
the control blocks it heads, that the priority map agrees with the lists,
and that each execute pointer is a member of the list its own priority
field names. Every walk is bounded, so a corrupted ring costs an assertion
and not a hang, and no test in the suite can hang. Expectations are counted
inside a scenario and gated between scenarios, so a failing kernel reports
how much it failed by without being driven further into its own
corruption.

The consequences are demonstrated from the terminator's context rather than
the completing thread's, which is what makes the pre-fix behaviour an
assertion instead of a wedged simulator. Compiled against the unfixed
sources the test fails 9 of the 24 expectations it reaches in the
uniprocessor tree and 10 of 24 in the SMP tree, and the failures are the
finding: the callback-side delete succeeds, the recreate succeeds, the
consistency snapshot disagrees, the target is neither terminal nor detached
when the service returns, and a peer has left the ready ring the recreated
block hijacked.

The SMP tree gets a second test for the case that needs concurrency. The
victim is excluded to core 1 and spins there without relinquishing while
the controller, excluded to core 0, terminates it, so the callback runs on
one core while the target executes on another. The callback-side reset and
delete must both be refused, and a sentinel written into the unused low end
of the victim's stack must survive -- a reset would have memset the whole
stack before rebuilding the frame. Every wait is bounded, and if the remote
precondition cannot be established the test says so and drops only the
assertions that depend on it rather than reporting a pass it did not earn;
measured over twenty consecutive runs it established the precondition every
time.

TX_NOT_INTERRUPTABLE and TX_DISABLE_ERROR_CHECKING are not among the five
build configurations either tree compiles, and each tree builds the whole
library once per configuration, so neither can be reached from inside the
suites. The lines this change adds under TX_NOT_INTERRUPTABLE are therefore
in no configuration the trees build, and they are where the permanent-
undeletability failure mode lives, so they get their own harness rather
than a compile check: the four sources plus the two error wrappers are
compiled directly into a test executable, once per combination, with
recorders standing behind the scheduler services they call. That is what
makes the marker's value at the moment of detachment directly observable.
It runs 66 expectations under TX_NOT_INTERRUPTABLE, 66 under that with
error checking disabled, 45 under that with notification disabled, and 63
under error checking disabled alone; against the unfixed sources those fail
25, 22, 6 and 20 respectively. The harness lives in the uniprocessor tree
only, because the four sources are identical between the kernels and the
shim replaces the very primitive the SMP port differs in, so a second copy
would compile the same text under the same macros. It is deliberately left
out of the coverage instrumentation, since the same source under different
feature macros has a different line set and merging those would confuse the
union rather than add to it.

Results. Both suites pass in all five configurations with GCC 14: 103 of
103 in the uniprocessor tree, up from 98, and 116 of 116 in the SMP tree,
up from 114. Merged line coverage is 100% in the uniprocessor tree and
5172 of 5183 in the SMP tree, whose eleven uncovered lines are the same
eleven that were uncovered before this change and are in tx_byte_pool_search
and tx_thread_smp_utilities; all four changed files are at 100% line
coverage in both trees, and SMP branch coverage rises from 2819 of 3548 to
2831 of 3556. Cross-compiled with arm-none-eabi-gcc at -Wall -Wextra for
Cortex-M4 against common and for Cortex-A7 SMP against common_smp, all four
files produce no diagnostics at all and an identical warning set to before
the change, under -std=gnu99 and -std=c99 alike -- unlike the module ports,
-std=c99 does not fail on these base ports, and even -Wconversion is clean.

No MISRA deviation is required: explicit comparisons to TX_TRUE, existing
ThreadX types, single-entry and single-exit control flow, no goto, and two
added constant-time tests that change no real-time complexity.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-10 12:04:40 -04:00
Frédéric Desbiens 9ee198a2d6 Made the stack analyze binary search require several consecutive fill words, so unwritten holes in a used stack no longer cause the highest stack pointer to be under-reported (#720)
The binary search in _tx_thread_stack_analyze() accepted a probe location as
unused as soon as a single word still held TX_STACK_FILL. A word inside an
otherwise used region that simply was never written - the padding of a
partially initialized local array, for example - therefore made the search
move away from the real boundary and report far less stack usage than the
thread had actually consumed, which in turn kept the stack guard from firing.

The probe now walks down from the candidate location and requires
TX_THREAD_STACK_ANALYZE_FILL_WORDS consecutive fill words before it treats
the location as unused, stopping early at the lowest location already known
to hold the fill pattern. The new macro defaults to eight words and can be
overridden in tx_port.h; setting it to one restores the previous behavior.
Holes shorter than the configured run no longer mislead the search, and the
result is unchanged for stacks that contain no such holes.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-10 11:02:42 -04:00
Frédéric Desbiens d56f96f37f Refreshed the reason gcc_check leaves RISC-V out (#722)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
r52_fvp / r52 (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The header explained the exclusion by saying RISC-V "is not regressing" and
that adding it would widen the toolchain download. The second half is still
true; the first read as though nothing in CI exercised the family at all,
which stopped being the case with #717.

RISC-V is now the best-covered of the four excluded families rather than the
least: regression_test.yml builds both ports and runs 955 tests on them under
QEMU -- 475 on RV32 and 480 on RV64, across five build configurations each --
which is more than a compile-and-link check could establish. That is a
stronger argument for leaving it out of this workflow than the original, so
the sentence now makes it.

The download figure is kept and quantified: the two bare-metal toolchains this
workflow would have to fetch are about 500 MB apiece.

Comment only; no behaviour change. scripts/check_gcc.sh names the same four
families but states the exclusion without giving a reason for it, so it needs
no matching edit.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-10 09:31:58 -04:00
Frédéric Desbiens 6c84e61d19 Gave install_riscv.sh the network hardening install.sh already had, and shared it between them (#721)
install.sh grew a retry loop, per-command timeouts and a deliberately
non-gating apt-get update after this runner pool cost several whole runs: a
mirror going silent for two hours, and a Hash Sum mismatch from a third-party
repository the project does not even use turning builds red. Those lessons were
local to that one file.

install_riscv.sh had none of them, and #717 puts it on every pull request's
critical path. Under set -e its bare apt-get update was a single point of
failure for the whole suite -- the precise case install.sh downgrades to a
warning on purpose -- and its two wget calls, each fetching about 500 MB, had
no retry and no timeout.

Rather than copy the helpers and let them drift again, they move to
tx_ci_common.sh and both scripts source it, following the arrangement
scripts/tx_windows_common.ps1 already uses on the Windows side. install.sh
keeps its behaviour exactly: same APT_OPTIONS, same 120-second TIMEOUT, same
three-attempt retry, and the comments explaining each of them travel with the
code they explain.

Two things are new:

  - TIMEOUT_LONG, 180 seconds, for a single large download. Sized against the
    39 seconds each tarball took on 10 Sep 2026 and deliberately not larger:
    the install step is capped at ten minutes, and a per-attempt timeout able
    to swallow that cap would leave the retry loop no turn to take, which is
    the failure mode the apt comment already records.

  - fetch(), which verifies a SHA-256 before anything is unpacked. Both digests
    were taken from the releases API and then checked against the bytes the CDN
    actually serves. This is not an independent trust root -- expected value and
    file come from the same host -- but it pins the bytes, so a deleted and
    re-pushed tag or a replaced asset stops the build instead of being picked up
    silently.

Also verifies qemu-system-riscv32 alongside riscv64. run.sh selects one per
architecture, so both are worth failing on here rather than at the first test.

Verified locally: retry returns 0 on success and 1 after three attempts;
fetch accepts a correct digest and, on a wrong one, fails and removes the
partial file; the source line resolves from the repository root, from an
absolute path and through a symlink; and both recorded digests match the
bytes served for the pinned tag.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-10 09:28:44 -04:00
Akif EjazandFrédéric Desbiens 18abe103af Enabled the RISC-V regression suite in CI for QEMU targets (#717)
* Enable the RISC-V regression suite in CI

The riscv: job in regression_test.yml has been present but commented out
since the RISC-V CI infrastructure landed (scripts/install_riscv.sh,
scripts/build_tx_riscv.sh, scripts/test_tx_riscv.sh, and the CMake tree
under test/tx/cmake/riscv/). It has been gated on a note that read
're-enable when RISC-V CI is ready', with no other blocker recorded.

The suite is ready. A local run of scripts/build_tx_riscv.sh followed by
scripts/test_tx_riscv.sh against upstream/dev builds every RISC-V target
and passes every registered test across all ten build configurations:

  RV32 (five configs): 5 * 95 = 475 tests, 0 failures, 0 timeouts.
  RV64 (five configs): 5 * 96 = 480 tests, 0 failures, 0 timeouts.

The extra RV64 test is threadx_riscv_new_thread_fpu_state_test, which
lives beside test/tx/cmake/riscv/regression/CMakeLists.txt and is added
only when THREADX_ARCH is risc-v64 -- the RV32 stack builder still
leaves the mstatus slot of the frame unwritten. Every test runs on
qemu-system-riscv32 / qemu-system-riscv64 with -machine virt, and each
configuration completes in roughly thirteen to fifteen seconds.

The job is wired the same way the ThreadX, SMP and FreeRTOS suites are:
it calls .github/workflows/regression_template.yml with the RISC-V
install/build/test scripts and the RISC-V cmake_path, sets
result_affix: RISC-V so its check name and artifacts are named, and
carries skip_deploy: true because coverage publishing stays on the
Linux suites for now. skip_coverage: true is kept because the CMake
configurations under test/tx/cmake/riscv match the Linux suite's minus
the coverage instrumentation, so gcovr has nothing to read -- the same
reason the FreeRTOS lane sets it. The regression_test.yml triggers were
not touched, so the RISC-V suite now runs on push and pull_request to
master and dev alongside the other three suites.

* Said why the RISC-V suite collects no coverage, instead of implying a missing build configuration

The comment added with the job read "No coverage build configuration for the
RISC-V suite yet". Both halves of what followed are true -- the configurations
under test/tx/cmake/riscv are the Linux suite's minus the coverage one, and
gcovr has nothing to read -- but "yet" points the next reader at a fix that
would not work.

Coverage here is not one missing default_build_coverage entry. These tests are
bare-metal images run under QEMU with -bios none, and
test/tx/cmake/riscv/bsp/syscalls.c has _write to the UART, _exit through the
sifive_test device, and stubs for _close, _fstat, _isatty, _lseek, _read and
_sbrk -- but no _open. gcov emits a .gcda by opening a path, so instrumenting
these builds produces nothing regardless of how they are configured. Nor is
the template's TX_COVERAGE=OFF what holds coverage off: test/tx/cmake/riscv is
its own top-level project and never declares that option, so the value is
inert there.

Getting a figure out of this suite needs a transport off the target -- gcov's
dump routines over the UART the BSP already drives, or semihosting. Worth
stating plainly in a project that asks for 100% coverage, rather than leaving
a reader to discover it after adding a configuration that cannot help.

Comment only; the job's inputs are unchanged.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

---------

Co-authored-by: Frédéric Desbiens <frederic.desbiens@eclipse-foundation.org>
2026-09-10 09:08:11 -04:00
Frédéric Desbiens 42f1d2bcf0 Pointed the SMP coverage comment at the pull request that made its measurements (#719)
The paragraph explaining why the SMP floor sits at 99 rather than 100 ended
with a bare reference that resolves to nothing a reader of this repository can
open. It named a source outside the tree in place of one inside it.

The real source is #677: it wrote the four tests that closed 53 of the 64
lines, took the measurements the paragraph quotes -- the 180,003 windows with
zero handovers, and the three remaining lines in tx_thread_smp_utilities.c --
and raised the floor from 98 to 99. Citing it gives the next reader somewhere
to go.

This is the same correction #718 made across five other files; this line was
missed because it sits in regression_test.yml rather than the template.

Comment only; no behaviour change.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-10 09:06:13 -04:00
Frédéric Desbiens 39277cf026 Stated the compiler and coverage requirements directly, and said who the pinned toolchain default serves (#718)
* Stated the compiler and coverage requirements instead of citing a file no contributor can open

Seven comments across five files cited a maintainer-local document as the
source for two project requirements: that GCC 14 on Linux is the default
compiler, and that the coverage target is 100%. That document is not part of
this repository and is not published anywhere, so the citation gave a reader
nothing to follow -- it named a source they cannot open, in place of simply
stating the requirement.

Both requirements are real and both stay. Only the pointer goes: each comment
now states the requirement on its own terms, which is what the surrounding
prose was already doing everywhere else.

  cmake/cortex_r52.cmake                     the pinned reference toolchain
  scripts/check_gcc.sh                       why the script exists
  .github/workflows/gcc_check.yml            why the workflow exists, and the
                                             GCC_VERSION pin
  .github/workflows/r52_fvp.yml              the GCC_VERSION pin
  .github/workflows/regression_template.yml  the coverage floor, twice

Comments only; no behaviour changes. Two paragraphs are rewrapped where the
shorter text left a ragged line. Verified that scripts/check_gcc.sh still
parses and prints its help from the header range it slices, and that
cmake/cortex_r52.cmake still configures the Cortex-R52 build.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Said who the pinned toolchain default in the CMake toolchain files serves

Both cmake/cortex_r52.cmake and cmake/cortex_m52.cmake default
ARM_TOOLCHAIN_PATH to a toolchains directory under the user's home, guarded by
an EXISTS check. Nothing said whether CI relies on that, and the natural
reading is that it does.

It does not. The three workflows that install a toolchain unpack it into the
workspace and cache it there, and r52_fvp.yml puts that directory on PATH
before configuring; scripts/check_gcc.sh passes -DARM_TOOLCHAIN_PATH at each of
its three CMake call sites. On a runner the guarded directory is absent, the
EXISTS check falls through, and the compiler comes from PATH. The default only
ever fires on a developer machine, where it is what makes a no-flag build work.

Both comments now say that, so the default is not mistaken for a CI dependency
and not removed as dead code. cortex_r52.cmake carries the explanation and
cortex_m52.cmake refers to it, matching the cross-reference already there.

The r52 comment also claimed absolute paths mean "the build does not depend on
PATH ordering", which is only true where the pinned directory exists -- in CI
the build depends on PATH and nothing else. Qualified accordingly.

Comments only; no behaviour changes. Verified that both toolchain files still
configure, and that the fall-through is real: with HOME pointed at a directory
holding no toolchains, cortex_r52.cmake configures against the arm-none-eabi-gcc
found on PATH.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-10 07:55:11 -04:00
Frédéric Desbiens 5e9d046b12 Compiled the module manager C sources with GCC, as clang already did (#716)
#689 added the module manager stage to check_clang.sh alone. The GCC half was
never written, so the module manager C stayed unbuilt by the project's declared
default compiler: 28 files of portable module manager under common_modules,
plus the three to nine per-port files under
ports_module/<core>/gnu/module_manager/src, across nine Arm module ports.

The stage is deliberately check_clang.sh's, port for port and header for
header, because a port covered by one check and not the other implies a parity
the checks list does not have. The same two details the ports dictate carry
over: an SMP port's control blocks come from common_smp rather than common, and
the TrustZone ports need -mcmse for their cmse_nonsecure_entry functions to be
honoured rather than ignored. The clang-only waiver does not: GCC implements
the optimize attribute that tx_thread_secure_stack.c carries, so nothing needs
suppressing for that file.

One divergence is forced by the toolchain. txm_module_manager_absolute_load.c
carries a #pragma message steering callers to the extended entry point, and the
C stages treat any compiler output as a failure. check_clang.sh silences it with
-Wno-#pragma-messages; GCC has no equivalent, and neither -Wno-pragmas nor any
other -W option suppresses the note -- verified with 14.3.rel1. The note is
therefore filtered out of the stage's output instead, together with the source
quote GCC prints beneath it. The filter stops at the next line that begins a
diagnostic of its own, so an error immediately following a waived note is still
reported; that case is what the injected-defect run below checks.

The workflow needed no trigger change: #689 added common_modules/** to both
path lists in advance, for the stage that had yet to arrive. Its header comment
is brought in line with what the workflow now runs.

Verified with the toolchain CI pins, arm-gnu-toolchain 14.3.rel1, both triples.
The full script passes and every port compiles every file, matching the counts
check_clang.sh reports for the same nine ports:

  cortex_a35 31/31   cortex_a35_smp 31/31   cortex_a7 34/34
  cortex_m0+ 33/33   cortex_m23 37/37       cortex_m3 33/33
  cortex_m33 37/37   cortex_m4 33/33        cortex_m7 33/33

Verified that the stage fails as intended by injecting defects into a throwaway
worktree: one in common_modules on the line straight after the waived pragma
note, reported under all nine ports, and one in a single port's own source,
reported only under that port.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-10 07:25:37 -04:00
Frédéric Desbiens 7f6e496233 Added the missing VFP enable field to the Cortex-R5/AC5 thread control block, which VFP builds were writing over the FileX pointer (#715)
The Cortex-R5/AC5 port defines TX_THREAD_EXTENSION_2 as empty while its
assembly reads and writes the per-thread VFP enable flag at [thread, #144].
With TX_ENABLE_VFP_SUPPORT that offset lands on tx_thread_filex_ptr, so
tx_thread_vfp_enable corrupted the FileX pointer and lazy save/restore
tested an unrelated value.

Defined TX_THREAD_EXTENSION_2 as ULONG tx_thread_vfp_enable, matching the
Armv7-A ports and the Cortex-R4/R5 GNU and AC6 ports fixed earlier, and
added tx_port_offset_check.c so a future layout change breaks the build
instead of silently retargeting the accesses. The offset was measured at
144 for this port and asserted.

Fixes #382

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-09 17:17:37 -04:00
tardigradeandFrédéric Desbiens a8c84593dd Added ARMv7-A SMP Linux build scripts(A5,A7,A9) (#674)
* ports: add Linux build scripts for Cortex-A5 SMP

* ports: fix Cortex-A5 SMP sample linking on Linux

* ports: add Linux build scripts for Cortex-A7 and A9 SMP

* ports: fix Cortex-A7 SMP assembly source path

* wip: add ATFE selection to Armv7 SMP scripts

* ports: complete Armv7 SMP Linux toolchain support

* ports: add local startup for Armv7 SMP samples

* ports: add license headers to Armv7 SMP scripts

* Dropped the preprocessing flag the file extension now carries

Two lines named assembly sources that #672 renamed. That change moved
twenty-nine files under gnu trees from .s to .S, because GAS runs the C
preprocessor on .S and not on .s: in a .s file every # line is a comment, so a
#define is never substituted and an #if/#else pair emits both arms. Four files
were silently doing the wrong thing as a result, including
ports_smp/cortex_a7_smp/gnu/src/tx_thread_smp_unprotect.s, which ignored all
four of its own feature macros.

These scripts had worked around the same defect with -x assembler-with-cpp
rather than hitting it, which was correct when they were written. With the
rename the flag is redundant and the lowercase names no longer resolve, so
cortex_a5_smp/build_threadx.sh failed with "cc1: fatal error:
tx_initialize_low_level.s: No such file or directory". The other seven
references to those two files across these three scripts already named them
with a capital S.

Verified with the pinned Arm GNU 14.3.rel1 rather than the 13.2 that a distro
package supplies: build_threadx.sh and build_threadx_sample.sh both succeed for
a5, a7 and a9, all three link a sample_threadx.out, and scripts/check_gcc.sh
passes end to end with its example stage reading 45 of 45, up from 42.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Frédéric Desbiens <frederic.desbiens@eclipse-foundation.org>
2026-09-09 16:30:23 -04:00
Frédéric Desbiens 383dd311cd Removed the vector table offset register and system stack pointer setup from the Cortex-M low-level initialization (#714)
_tx_initialize_low_level wrote VTOR and derived _tx_thread_system_stack_ptr from
the reset vector on every Cortex-M port. Programming VTOR is the job of the
low-level startup code that runs before the kernel is entered, and doing it again
inside ThreadX silently overrode a vector table that the application, a
bootloader or a firmware update had already installed. It also forced every
application to export a vector table symbol that the kernel itself never needed.

_tx_thread_system_stack_ptr is never read by any Cortex-M port, so the value
copied out of the reset vector served no purpose.

Both blocks are removed from all Cortex-M ports, together with the now-unused
symbol declarations. The GreenHills files keep the NVIC base address load that
the removed block used to leave in r0 for the SysTick setup that follows.

Applications that relied on ThreadX programming VTOR must now set it in their
startup code. The bundled examples link their vector table at address zero and
run on the reset default.

Fixes #370

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-09 16:28:59 -04:00
850a172bac Added lazy FPU stacking and QEMU functional tests for RV64 GNU port (#549)
* add lazy FPU stacking to context save/restore

Save mstatus/sstatus to stack slot 29 and skip floating-point register
save/restore when FS is Off (bits 14:13). This avoids unnecessary FP
context work for threads that do not use the FPU.

- context_save: check FS in nested and first-level interrupt paths
- context_restore: gate FP restore on nested, no-preempt, and preempt paths
- use sstatus when TX_RISCV_SMODE is defined, otherwise mstatus

* add QEMU virt CMake build and automated test runner
Wire the QEMU virt demo into the CMake build system and add a
Python/GDB functional test runner, mirroring the risc-v32/gnu port.
- Add qemu_virt/CMakeLists.txt to build kernel.elf and register the
  check-functional-riscv64 target (requires Python3; skipped if absent)
- Link kernel.elf with --whole-archive so all ThreadX symbols resolve
- Pin _start at 0x80000000 via .text.boot in entry.s and
  KEEP(*(.text.boot)) in link.lds
- Extend demo_threadx.c with fpu_test_val and shorten thread_0 sleep
  for GDB-driven FPU, timer, and preemption checks
- Add test/azrtos_test_tx_gnu_riscv64_qemu.py; verified passing on
  QEMU virt (FPU, timer interrupt, preemption)

* Clean up RV64 PR scope and remove QEMU test integration leftovers

Revert accidental RV64 qemu_virt test/CMake integration changes and keep this branch
focused on lazy FPU context handling only. Also remove unintended TX_RISCV_SMODE-based
mstatus/sstatus save path and align comments/logic to mstatus-only behavior.

* Initialize mstatus.FS in RV64 stack build so new threads start with clean FP state

Slot 29 was left uninitialized while context restore reads it as an FP-live
hint; garbage FS bits could make a new thread inherit the previous thread's
floating-point registers.

* Add RV64 regression test for the FP state of a newly created thread

The test dirties every floating point register, then creates a thread and
checks that the stack builder wrote the mstatus slot and that the new
thread starts with all floating point registers zeroed. It is registered
for RV64 only, since the RV32 stack builder still leaves the slot unwritten.

* Completed the RISC-V64 lazy FPU so the restore side matches the save side

The lazy FPU save in this branch skips the floating-point stores when
mstatus.FS is Off, and records the mstatus it judged that on in frame slot
29. Merged onto current dev, only the save side had that treatment: both
restore paths and the scheduler's interrupt-frame path still reloaded the
FP registers unconditionally, from slots the save had deliberately left
alone. A thread that never touched the FP unit would have had whatever the
frame happened to contain loaded into its registers, and FS driven to
Dirty on the way out.

The guard is added at the three places that consume an interrupt frame:
both paths in _tx_thread_context_restore, and _tx_thread_schedule_loop.
Each reads slot 29 and skips the FP block when FS was Off, which is the
same shape the risc-v32 port already uses.

The solicited path is deliberately left alone. _tx_thread_system_return
saves the callee-saved FP registers unconditionally, so restoring them
unconditionally is consistent; making that pair lazy as well is a separate
change, and risc-v32 is the model for it.

Verified with QEMU on all five configurations:

  risc-v64   96 of 96 passing, five configurations, nothing unlinkable
  risc-v32   95 of 95 passing, five configurations, unchanged
  functional check-functional-riscv64 passes every check

The ninety-sixth test is the one this branch adds. It is load bearing:
seeding stack build with FS = Off instead of Initial makes it fail, and
restoring the seed makes it pass, so it guards the behaviour the rest of
this branch is about rather than passing regardless.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

---------

Co-authored-by: r <r@r>
Co-authored-by: Frédéric Desbiens <frederic.desbiens@eclipse-foundation.org>
2026-09-09 16:25:28 -04:00
Frédéric Desbiensandr 16297d7a2a Added a workflow that runs the Cortex-R52 images on the Armv8-R AEM FVP (#688)
* Added a workflow that runs the Cortex-R52 images on the Armv8-R AEM FVP

Nothing in CI executed a single instruction of any ThreadX port. gcc_check
compiles and links -- its own header says it "executes nothing" -- and its
CMake stage covers one R52 configuration, the base FVP example. A change that
assembles cleanly, links cleanly and then hangs on the first context switch
passed every required check.

This workflow builds both supported configurations and runs them on the
Armv8-R AEM FVP, asserting each image's self-reported result. The feature
configuration matters as much as the default one: VFP with the hard float ABI,
FIQ, IRQ nesting and FIQ nesting gate whole assembly blocks that the default
configuration never assembles, and it builds eight images where the default
builds five.

Arm distributes the model free of charge but behind a click-through licence
with no stable unauthenticated URL -- the developer.arm.com permalink forms
all 404 for it -- so the download location is a repository variable rather
than a literal. When it is unset the build lanes still run and still gate the
pull request; only execution is skipped, and it says so in the log and in the
job summary rather than passing quietly.

The image list is read from the generated ninja graph rather than kept in the
workflow. The images are EXCLUDE_FROM_ALL, so a bare build reports "no work to
do", and reading the graph means a target added to CMakeLists.txt cannot
escape the check by nobody remembering to list it here.

The module manager port and the S32Z280 targets are not on dev yet. Their
lanes belong in this workflow when they land, not in a second one.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Kept the FVP cache honest and stopped ctest passing on an empty suite

Two corrections to the new workflow.

The FVP cache key named the URL and hashed the workflow file. hashFiles()
reads files, so it cannot hash a repository variable, and pointing it at
this workflow keyed the cache on the file's own contents instead. That is
wrong in both directions: repointing FVP_AEMV8R_URL at a different model
kept serving the old one from cache, which is exactly what the comment
claimed the key prevented, and editing anything in this file forced a
needless refetch of a large archive. A small step hashes the URL and the
key uses that, so it now tracks what it names.

ctest reported success when it found nothing to run. Its default is to
pass an empty suite, and the example registers its tests inside both
if(FVP found) and if(Python3_FOUND). Either one going unmet on a runner
would have left this stage green having executed no image at all, which
is the hole the workflow exists to close. --no-tests=error makes an empty
suite a failure.

Verified with the pinned toolchain, Arm GNU 14.3.rel1: both configurations
configure, enumerate and build, and the image lists are the ones claimed.

  default: 5 images  boot_check demo_m2 demo_m3 demo_mpu demo_threadx
  feature: 8 images  the above plus demo_fiq demo_m5 demo_nesting

Also checked that the target enumeration survives the real ninja graph:
the generated graph offers twenty .elf paths, and the filter reduces them
to those eight, dropping the cmake_object_order_depends_target aliases and
the CMakeFiles copies.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

---------

Co-authored-by: r <r@r>
2026-09-09 14:39:54 -04:00
Frédéric Desbiensandr 8be4c45cce Compiled the module manager C sources, which no check had ever built (#689)
* Compiled the module manager C sources, which no check had ever built

#672 corrected this script's assembly glob and brought the module ports into
the count, but only their assembly. Their C stayed outside every check: 27
files of portable module manager under common_modules, plus the per-port code
under ports_module/<core>/gnu/module_manager/src. By this script's own standard
-- a port simply absent from the count reads as covered -- 293 files across
nine Arm module ports were compiled by nothing, with either compiler.

Each module port ships its own tx_port.h and txm_module_port.h carrying the
control-block extensions the dispatch code needs, so a port is compiled against
its own headers rather than the base port's. Two details the ports themselves
dictate:

  An SMP port's control blocks come from common_smp. Pairing cortex_a35_smp
  with the single-core headers hid _tx_thread_smp_protect and
  _tx_thread_smp_unprotect behind implicit declarations and lost
  tx_thread_smp_core_executing from TX_THREAD, so fourteen files reported
  errors for a port that builds correctly.

  The TrustZone ports carry cmse_nonsecure_entry, which needs -mcmse to be
  honoured rather than ignored. tx_thread_secure_stack.c also carries GCC's
  optimize attribute, which clang does not implement; that divergence is
  suppressed by name, for a file GCC builds cleanly.

All nine ports compile: 293 of 293. Verified that the stage fails as intended
by injecting a defect into a throwaway worktree -- a defect in common_modules
is reported under every port, one in a port's own source only under that port.

Assisted-by: Claude Code (Opus 5)

* Ran the module manager stage against the sources that reach it

Three corrections to the new stage, all found by running it against dev
rather than against the tree it was written on.

The workflows did not trigger on common_modules. Both check_clang.yml and
check_gcc.yml list ports_module but not common_modules, and the portable
module manager under it is the larger half of what the stage compiles --
28 of the 31 to 37 files each port builds, plus two include directories.
A change there left the stage unrun, which is the same "absent from the
count reads as covered" that the stage exists to close. Both lists gain
common_modules, and they stay identical to each other as the comment in
each asks.

A deliberate deprecation notice read as a build failure. Since this PR
was opened, txm_module_manager_absolute_load.c gained a #pragma message
steering callers to the extended entry point. The stage treats any
compiler output as a failure, so that one notice failed every port: nine
failures on a tree where nothing is wrong. Pragma messages are now
waived for the stage, because a notice to callers is not a defect in the
file that carries it.

That failure also printed nothing. Both C stages report by grepping the
output for "error:", so a diagnostic that is not an error produced a bare
FAIL line with no reason under it, and the only way to learn the reason
was to reproduce the compile by hand. Both stages now fall back to
showing what the compiler actually said.

The TrustZone attribute waiver is narrowed to the one file that needs it.
tx_thread_secure_stack.c carries GCC's optimize attribute, which clang
does not implement; it is the only file among the 300-odd this stage
compiles that does. Waiving the warning for the whole port would have
swallowed a stray unknown attribute anywhere else in it.

Verified with the same toolchain CI uses, ATfE 22.1.0: the full script
passes, and every port compiles every file.

  cortex_a35 31/31   cortex_a35_smp 31/31   cortex_a7 34/34
  cortex_m0+ 33/33   cortex_m23 37/37       cortex_m3 33/33
  cortex_m33 37/37   cortex_m4 33/33        cortex_m7 33/33

The counts are each one higher than this PR first reported, because
txm_module_manager_absolute_load_extended.c has landed since. The stage
picked it up with no edit, which is what globbing the directories was for.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

---------

Co-authored-by: r <r@r>
2026-09-09 14:39:30 -04:00
Frédéric Desbiens 6ce8d5cc76 Marked every published ThreadX include directory as SYSTEM so applications no longer get warnings from ThreadX headers (#713)
* Marked every published ThreadX include directory as SYSTEM so applications no longer get warnings from ThreadX headers

Commit 8c3c08f added the SYSTEM keyword to the target_include_directories
call in common/CMakeLists.txt, but the same call in every port, in the
SMP common directory's consumers, in the POSIX and FreeRTOS compatibility
layers, and for the generated tx_user.h directory was left unchanged.
A consumer building with a strict warning set therefore still saw
diagnostics coming from tx_port.h and from the compatibility layer
headers, which is exactly what the original change set out to avoid.

Added the SYSTEM keyword to all of those calls so CMake emits -isystem
rather than -I for every directory holding a ThreadX public header. The
ARMv7-M, ARMv8-M, ARMv7-A and ARMv8-A architecture sources under
ports_arch were updated alongside the ports they generate, keeping the
two in step.

Directories that are PRIVATE to an example or test build were left as
they are, since nothing is published from them.

Fixes #290

Assisted-by: Copilot (Opus 5) <noreply@github.com>

* Stopped apt-get update being a gate it was never meant to be

apt-get update fails if any configured repository serves a bad index,
including ones this project never reads. The GitHub runner image carries
Google's and Microsoft's apt repositories, and a Hash Sum mismatch from
Google's, their CDN caught mid-publish with the index and the Release
file eight hours apart, failed all three attempts and turned a run red
over a browser nobody was installing.

Made a failed update warn and carry on, leaving apt-get install as the
gate. Nothing is weakened by that: the install still exits on a package
it cannot find, so an unreachable archive still stops the script, one
step later and naming the package it could not get, which is a better
diagnostic than a hash mismatch in a repository nobody asked for.

Disabling third-party sources before updating would keep the update
strict, but this script also runs on a contributor's own machine, and
rewriting someone's apt configuration to suit CI would be worse than
tolerating a stale index for an archive we do not read.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-09 14:27:02 -04:00
4d90a0c21c Stopped the win32 and win64 ports from enabling performance metrics and event trace (#676)
* win32: do not always enable trace or performance metrics in tx_port.h

if these are required they can be enabled in tx_user.h

* win64: do not always enable performance metrics in tx_port.h

if required they can be enabled in tx_user.h

* Removed the disabled blocks rather than commenting them out

The win32 and win64 ports were the only two that turned performance
metrics on for the application, and win32 the only one that turned event
trace on. Leaving those to tx_user.h is right: the symbols extend the
control blocks, so a port that sets them behind the application's back
changes structures the application also sees.

The blocks were disabled with #if 0 rather than deleted. That is the form
MISRA C:2012 Directive 4.4 is about -- sections of code should not be
commented out -- and it leaves two copies of a list that now has no
reader. They are removed, and a short note in their place says where the
symbols belong and why the port does not set them.

No behaviour change beyond what this pull request already made. Checked
that the preprocessor nesting in both headers is still balanced.

Worth recording for whoever looks next: with this in, no port defines
either symbol. The linux port carries the same list commented out, which
reads at a glance like a third case but is not one.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

---------

Co-authored-by: r <r@r>
Co-authored-by: Frédéric Desbiens <frederic.desbiens@eclipse-foundation.org>
2026-09-09 13:39:13 -04:00
Akif EjazandFrédéric Desbiens 40db27e843 riscv32: spec compliance and regression test fix (#691)
* riscv32: spec compliance and regression test fix

Signed-off-by: Akif Ejaz <akifejaz40@gmail.com>

* Derived the RISC-V32 frame sizes from the port contract in one place

tx_port.h published TX_RISCV_TRAP_FRAME_SIZE for the GNU BSP assembly, but
nothing in the port consumed it. Six .S files each rebuilt the same numbers
from their own #if, so the interrupt frame size was written out in seven
places and the solicited frame size in three.

That is the shape that produced the RISC-V64 fault fixed in #708, where the
port moved to a padded frame and one copy of the constant did not. The
sources now include tx_port.h and take both sizes from it, and no literal
frame size remains in the port. TX_RISCV_SOL_FRAME_SIZE joins the contract,
since the solicited frame was never published at all.

The emitted code is unchanged: 400 and 176 bytes for ILP32D, 128 for
soft-float, confirmed by disassembly before and after.

Two further corrections:

_tx_initialize_low_level carried .global immediately followed by .weak, so
the symbol stayed weak and the .global did nothing. Weak is what the port
wants, because the example and regression BSPs both provide their own
definition, so the stray .global is removed rather than the .weak. Verified
with nm that the symbol is still W.

The QEMU runner seeded fpu_verified from skip_fpu, so a soft-float run
satisfied the FPU gate whether or not the script ever reported the skip. It
now starts false and is set only when the skip marker is present, so a run
that dies before reaching that point fails instead of passing.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

---------

Signed-off-by: Akif Ejaz <akifejaz40@gmail.com>
Co-authored-by: Frédéric Desbiens <frederic.desbiens@eclipse-foundation.org>
2026-09-09 11:40:05 -04:00
Frédéric Desbiens 3fc28d8979 Allowed a BASEPRI-masked interrupt to wake the Cortex-M idle loop from WFI (#711)
The idle loop in _tx_thread_schedule masks interrupts before re-reading
_tx_thread_execute_ptr, so that an interrupt cannot make a thread ready
between that read and the WFI and then be lost. With PRIMASK that is safe:
the architecture excludes PRIMASK from the WFI wake-up condition, so the
interrupt still wakes the processor and is taken as soon as the loop
re-enables interrupts.

BASEPRI is not excluded. When TX_PORT_USE_BASEPRI is defined, the mask
written at the top of the loop is still in place at the WFI, so any
interrupt at or below TX_PORT_BASEPRI fails to wake the processor at all.
On a system where every interrupt is managed by ThreadX, that is every
interrupt, and the processor stays in WFI or in the low power mode entered
by tx_low_power_enter until something outside the mask happens.

Set PRIMASK and clear BASEPRI for the duration of the WFI, then restore
BASEPRI and clear PRIMASK. Interrupts remain masked across the whole
window, so the original race is still closed, but the wake-up condition is
now evaluated with BASEPRI clear and any enabled interrupt can end the wait.
A pending interrupt is taken at the CPSIE i, or at the existing unmask
further down the loop if it falls below TX_PORT_BASEPRI.

The change is made in the ARMv7-M and ARMv8-M sources under ports_arch,
including the module manager, and propagated to the generated ports with
the copy scripts. The ghs and keil ports do not implement
TX_PORT_USE_BASEPRI, and Cortex-M0/M0+/M23 have no BASEPRI register, so
those ports are unaffected.

Verified by assembling every patched GNU port for Cortex-M3, M4, M7, M33,
M55 and M85, with and without TX_PORT_USE_BASEPRI, TX_LOW_POWER and
TX_ENABLE_EXECUTION_CHANGE_NOTIFY, and by checking the disassembly of the
idle loop.

Fixes #279

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-09 11:29:14 -04:00
Frédéric Desbiens 146d57b235 Fixed the RISC-V64 trap frame size mismatch in the regression test BSP (#708)
The RISC-V64 port moved its interrupt frame to 528 bytes (65 slots plus 8
bytes of padding, so sp stays 16-byte aligned at a call) and published the
size as TX_RISCV_TRAP_FRAME_SIZE. The port sources were converted to use
it, but the shared regression test BSP was not: its trap_entry still
allocated a hardcoded 65 * REGBYTES, or 520 bytes.

Every interrupt therefore unwound 8 bytes more than it allocated:

    trap_entry:                   addi  sp,sp,-520
    _tx_thread_context_restore:   addi  sp,sp,528

On the RISC-V64 regression suite that left 24 of 95 tests failing in the
default configuration, typically as an illegal instruction once execution
reached a corrupted frame. The example BSP under the port directory was
converted with the port and was unaffected, which is why the functional
QEMU test kept passing.

The test BSP now takes both frame sizes from the port it is linked
against, so the two cannot drift apart again. A port that publishes no
contract keeps the historical layout, so the RISC-V32 side is unchanged
until its own port publishes one.

TX_RISCV_TRAP_CALL_FRAME_SIZE is restored to the RISC-V64 tx_port.h. It
was removed as unused when the frame sizes were introduced, but it is
part of the same contract: it is the space a trap entry reserves around a
call into C, and the psABI requires 16 bytes there rather than one
register slot.

Verified on QEMU with every linkable test built, comparing against the
commit before the port change:

    before the port change   2 failures out of 95 (both unlinkable)
    current dev              24 failures out of 95
    with this change          2 failures out of 95 (both unlinkable)

The two remaining failures predate all of this: newlib pulls _impure_ptr
out of R_RISCV_HI20 range for time(), so those two binaries do not link.
RISC-V32 is unchanged at 2 failures across all five configurations.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-09 10:56:57 -04:00
Frédéric Desbiens 9a03838381 Stopped the test runners from reporting an incomplete build as test failures (#709)
The cmake test runners drove Ninja with its default keep-going of 1, so the
first failing target ended the build. Every target scheduled after it was
simply absent, and ctest reports a missing binary as a failing test. A
single link error therefore produced a failure count that moved with build
scheduling order rather than with the code.

Measured on the RISC-V64 regression suite, where two targets genuinely
cannot link:

    before   79 of 95 test binaries built
    after    93 of 95 test binaries built

Fourteen perfectly good binaries were being skipped and counted as
failures. Passing -k 0 lets Ninja finish everything it can; the build still
exits non-zero when a target fails.

Three related problems in the same paths are fixed with it.

A failing configuration used to abort the loop over configurations, so
under set -e the ones after it went unbuilt or untested. The build loops
and the serial test loops now accumulate status and return it at the end,
which is what the parallel test branch already did with wait, and what
cmake_bootstrap.sh already documented for ctest.

Capturing that status removes the set -e protection inside the functions,
so two latent faults become reachable and are closed here. A failed pushd
would have let ctest run in the source tree, where it finds no tests and
reports success; the pushd is now guarded. And ctest's status was
discarded by the popd that follows it, so a configuration with failing
tests returned 0 and was reported as a pass; the status is now carried
past the popd and the summary steps.

Verified on the RISC-V32 suite, which has two genuine failures in each of
its five configurations. Both the serial and the parallel branch now test
all five and exit 8, where the serial branch previously stopped after the
first configuration.

The tx and smp runners are symlinks to scripts/cmake_bootstrap.sh, so they
are covered by the one change there.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-09 10:56:28 -04:00
Frédéric Desbiens 37f48fa83c Removed FIFO queueing from the ARMv7-A SMP ports so an ISR can no longer deadlock waiting for protection (#707)
The Cortex-A5, A7 and A9 SMP ports guarded the inter-core protection with
a FIFO wait list: a core that could not take the lock added itself to a
queue, incremented _tx_thread_smp_protect_wait_counts[core], and only the
core at the head of the queue was allowed to acquire.

Two parts of that scheme require the waiting core to be interruptible.
_tx_thread_smp_unprotect() refuses to release the protection while the
releasing core's own wait count is non-zero, and a queued core is taken
back out of the list by _tx_thread_context_restore() when it is preempted.
Neither can happen on a core that is spinning inside an ISR with
interrupts already masked, because the spin loop restores the caller's
interrupt posture rather than enabling interrupts. The core stays in the
list forever and the system deadlocks with cores stuck in the wait loop.

This is the deadlock Microsoft removed from the ARMv8-A SMP ports in
6.1.11, by dropping the wait list and using a plain LDAXR/STXR spinlock.
The same removal was announced for the ARMv7-A ports at the time but was
never made, so those three ports have carried the deadlock ever since.

The FIFO queueing is now removed from the ARMv7-A SMP ports as well.
_tx_thread_smp_protect() becomes an LDREX/STREX spinlock that releases
interrupts between attempts, matching the sequence already used by the
Cortex-R8 SMP port; _tx_thread_smp_unprotect() no longer consults the
wait counts; and _tx_thread_context_restore() no longer has to unqueue a
preempted core. The now unreferenced wait list macro headers are deleted.

Verified by assembling all three GNU ports with arm-none-eabi-gcc for
cortex-a5, cortex-a7 and cortex-a9, with and without
TX_ENABLE_FIQ_SUPPORT, TX_ENABLE_WFE and TX_MPCORE_DEBUG_ENABLE, and by
reading back the disassembly of the new protect sequence. The port
consistency checks pass.

Fixes #219

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-09 10:29:51 -04:00
Frédéric Desbiens d3fb72b6dd Aligned the simulator ports' fake stack pointer so ThreadX no longer performs misaligned ULONG accesses (#705)
The Linux, Win32 and Win64 simulation ports build a fake initial stack
pointer by subtracting a fixed 8 bytes from tx_thread_stack_end. That
field addresses the last byte of the thread's stack area, so it is one
less than an aligned address and the resulting pointer is misaligned by
construction, no matter how well aligned the stack the application
supplied was.

Two ULONG accesses then use that pointer. _tx_thread_stack_build() itself
clears the word below it, and _tx_thread_create() copies it into
tx_thread_stack_highest_ptr, which TX_THREAD_STACK_CHECK dereferences on
every suspend and resume when TX_ENABLE_STACK_CHECKING is defined. Both
are undefined behaviour. They happen to work on x86 but are reported by
GCC's undefined behaviour sanitizer, and would fault on a host that
requires natural alignment.

The fake stack pointer is now rounded down to a ULONG boundary, which
leaves it inside the stack area and makes both accesses aligned.

Verified by building the Linux port with -fsanitize=undefined and
running the demo: the two reported diagnostics are produced before the
change and neither appears after it. The tx and smp regression suites
pass, 98 and 114 tests respectively.

Fixes #218

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-09 09:34:33 -04:00
Frédéric Desbiens 2e48ce0d1b Removed the stray build artifacts committed with the RISC-V64 port fix (#706)
A local CMake build tree was staged by mistake alongside the RISC-V64
spec compliance work in #698, putting nine generated files on dev,
including a compiled kernel.elf, build.ninja, the Ninja dependency logs
and a QEMU run log. None of it belongs in the repository.

The root .gitignore listed build directories by name rather than by
pattern, so build/, build_qemu/, build_m7/ and the build_r52 variants
were covered but a differently named tree was not. Those entries are
replaced with a single build*/ pattern, which covers every existing name
and any future one. No tracked file matches the new pattern.

The artifacts remain reachable in history; only the working tree is
corrected, since rewriting a shared branch is the greater harm.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-09 09:31:00 -04:00
Akif EjazandFrédéric Desbiens 164f211a01 riscv64: spec compliance and regression test fix (#698)
* spec compliance

Signed-off-by: Akif Ejaz <akifejaz40@gmail.com>

* revert the demo changes

Signed-off-by: Akif Ejaz <akifejaz40@gmail.com>

* Restored the FPU demo hook so the RISC-V64 functional test can pass

The new check-functional-riscv64 target verifies FPU context switching by
watching fpu_test_val advance by 1.1f on each pass through
thread_6_and_7_entry. The GDB script deliberately treats a missing symbol
as a failure rather than silently skipping the check, but the demo no
longer defined it, so the target failed on every run:

    FPU_VERIFIED_FAIL_NO_SYMBOL

The definition and the increment are restored, matching what the risc-v32
demo already carries. The functional target now passes end to end.

Three small corrections are folded in:

- tx_port.h carried a comment stating that the ISA string must include
  Zicsr, but nothing enforced it, so an rv64imac build failed with a wall
  of assembler "unrecognized opcode" errors. It now stops at one clear
  diagnostic.
- Removed TX_RISCV_TRAP_CALL_FRAME_SIZE, which nothing referenced.
- The example .gitignore listed qemu-riscv32.log, but the runner writes
  qemu-riscv64.log, so the generated log showed up as an untracked file.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

---------

Signed-off-by: Akif Ejaz <akifejaz40@gmail.com>
Co-authored-by: Frédéric Desbiens <frederic.desbiens@eclipse-foundation.org>
2026-09-09 09:26:11 -04:00
Frédéric Desbiens 515ab8aba1 Added the missing memory barrier so ARMv7-A SMP schedulers no longer miss preemptions (#704)
_tx_thread_schedule stores the newly selected thread into
_tx_thread_current_ptr[core] and then reloads _tx_thread_execute_ptr[core]
to detect a concurrent scheduling decision made by another core. On the
other side, _tx_thread_smp_core_interrupt stores the execute pointer and
then reads the current pointer to decide whether an inter-core interrupt is
required. This is a store-buffer pattern: without a barrier both cores can
observe stale values, the interrupt is skipped, and a ready thread with the
highest priority is never scheduled.

The barrier was added to the ARMv8-A SMP scheduler in 6.2.1, but the
ARMv7-A SMP ports were left untouched even though they implement the same
protocol. This adds the corresponding DMB to the Cortex-A5, Cortex-A7 and
Cortex-A9 SMP schedulers for both the AC5 and GNU toolchains.

The GNU variants were verified by assembling them with arm-none-eabi-gcc
13.2.1 for their respective cores.

Refs #209

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-09 09:10:36 -04:00
Frédéric Desbiens e99f0d4207 Corrected the module kernel stack size so it no longer overstates the usable stack (#701)
The module manager recorded tx_thread_module_kernel_stack_size as the raw
TXM_MODULE_KERNEL_STACK_SIZE constant, but the end of the kernel stack is aligned
downwards to an eight-byte boundary while _txm_module_manager_object_allocate only
guarantees ULONG alignment. The recorded size could therefore overstate the usable
stack by up to seven bytes. The scheduler copies this value into tx_thread_stack_size
whenever a user mode module thread enters the kernel, so the overstated value is
visible to RTOS-aware debuggers and to anything built on it.

The size is now derived from the aligned end minus the start. Also documented that
TX_ENABLE_STACK_CHECKING is not supported for module threads.

Refs #181

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-08 16:44:48 -04:00
Frédéric Desbiens 945f5f5caa Moved the ARC ISR enter callout onto the system stack (#700)
_tx_thread_context_save() calls _tx_execution_isr_enter when
TX_ENABLE_EXECUTION_CHANGE_NOTIFY is defined. On the path where an interrupt
preempted a running thread, that call was made before the switch to the system
stack, so the 32-byte call frame and the whole stack footprint of the callout
were taken from the interrupted thread's stack, on top of the 160-byte interrupt
frame the port had already allocated there.

The callout is supplied by the application, so its stack usage is not bounded by
ThreadX, and it is charged to every thread that happens to be running when an
interrupt arrives.

The two other callout sites in the same routine, the nested save and the idle
system save, already run on the system stack, as does the _tx_execution_isr_exit
call in _tx_thread_context_restore. _tx_thread_schedule was reordered in 6.1.9
so that _tx_execution_thread_enter runs on the system stack rather than the
thread stack; the same reorder was never applied to the context save.

The switch to the system stack now happens before the callout in the ARCv2_EM,
ARC_HS and SMP ARC_HS ports. _tx_thread_context_fast_save is unchanged because
the fast interrupt path never switches stacks by design.

Fixes #149

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-08 15:42:39 -04:00
tardigradeandFrédéric Desbiens 29afcc3946 Fixed a kernel stack leak when deleting user-mode module threads (#692)
* modules: free kernel stack on thread deletion

Signed-off-by: Prashit Vora <prashitvora2006@gmail.com>

* Preserved the thread object release when the kernel stack cannot be freed

Releasing the kernel stack ahead of the thread object made a failure of the kernel
stack deallocation abort the thread object release. The thread had already been
deleted at that point, so the thread object would have stayed allocated for the
lifetime of the module.

The thread object is now always released once the delete succeeds, and the kernel
stack failure is reported only when it does not mask a thread object failure.



---------

Signed-off-by: Prashit Vora <prashitvora2006@gmail.com>
Co-authored-by: Frédéric Desbiens <frederic.desbiens@eclipse-foundation.org>
Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-08 11:17:57 -04:00
Frédéric Desbiens b0ec8bfbb9 Fixed the invalid module data pointers in the absolute module load (#699)
An absolutely located module has its code and its data placed at two
independent fixed addresses by the module's linker script. The module
preamble carries the code and data sizes but not the data address, so
_txm_module_manager_absolute_load() could not determine where the
module's data area was. It computed txm_module_instance_data_start
from the code size and the preamble size, which yields a size rather
than an address, and it set txm_module_instance_module_data_base_address
one past the end of the byte pool allocation.

Added _txm_module_manager_absolute_load_extended(), which accepts the
module's data area address from the caller. Deprecated
_txm_module_manager_absolute_load(), which now forwards to the extended
service with an unknown data area location and rejects modules that
request memory protection, since the memory protection hardware cannot
be programmed to cover an unknown data area.

Fixes #450

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-08 10:03:45 -04:00