Commit Graph
180 Commits
Author SHA1 Message Date
Frédéric Desbiens 70a5300977 Added a ThreadX module manager port for the Cortex-R52 (#639)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
r52_fvp / r52 (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
Added the first GNU ThreadX module port for an Arm R-profile core.

  The port combines the PMSAv8-R MPU model from Cortex-M33 with the AArch32
  privilege and processor-mode handling from Cortex-R4. It provides the complete
  module path: headers, manager sources, scheduler integration, user-mode entry,
  SVC dispatch, data- and prefetch-abort capture, fault notification, relocatable
  module loading, and shared-memory regions.

  Module isolation uses eight MPU regions (8–15) with a 64-byte granule. The
  scheduler replaces those regions on each module switch and manages a separate
  privileged loading window in region 16. Assembly-visible structure offsets and
  region-layout assumptions are checked at build time.

  Added independent module demonstrations for the S32Z280-594EVB and Armv8-R AEM
  FVP. The automated FVP regressions cover:

  - Loading the same position-independent module at different addresses
  - User-mode data and instruction access violations
  - Fault capture and notification for both abort types
  - Shared-region access, alignment, exhaustion, empty-size, and overflow handling
  - Required module-property combinations
  - The GCC CLZ-based priority search

  The port requires user mode and memory protection together, at least 17 EL1 MPU
  regions, and currently validates A32 modules; Thumb module execution remains
  unvalidated.

  Validated with GNU Arm 14.3.1. The FVP suites pass 8/8 in the default
  configuration and 11/11 with hard-float, FIQ, and interrupt nesting enabled.
  The S32Z280 images build cleanly, and the module isolation and relocation paths
  were exercised on S32Z280 silicon during development.

  Matching user documentation is provided by rtos-docs-asciidoc PR #41.

  Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-17 15:39:30 -04:00
Frédéric Desbiens ad558a7b1f Fixed the zero trace time stamps in the Linux ports' MISRA builds (#749)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
Both Linux ports define TX_TRACE_TIME_SOURCE as _tx_misra_time_stamp_get() when
TX_MISRA_ENABLE is set, and neither implements that function, so both inherit
the generic `return(0);` from tx_misra.c. Every trace event is stamped zero.
The buffer carries no timing, and the kernel's own check for an entry having
been overwritten -- time_stamp against the entry's own stamp, in the block and
byte allocates and in the system suspend and resume -- compares zero with zero,
so it never fires and a service patches whatever now occupies the slot.

Both ports now read in MISRA builds the clock they already read otherwise,
_tx_linux_time_stamp.tv_nsec, which TX_TRACE_PORT_EXTENSION refreshes on every
recorded event in both forms of the insert. The non-SMP port's non-MISRA macro
carried a trailing semicolon, which made it a statement and is why the MISRA
insert -- which takes the time source as a function argument -- could not use
it; that is dropped and the two branches become one definition. Both headers
keep the _tx_misra_time_stamp_get declaration, because tx_misra.c still defines
it and is compiled for these ports.

The MISRA insert evaluates its time source before the callee refreshes the
clock, so each entry carries the reading taken at the previous recorded event.
Stamps are real, distinct and ordered, which is what the overwrite check needs.

The trace entry update test gains an assertion that the buffer holds an entry
the port actually stamped. It fails on dev with ERROR #13 under
misra_trace_build and passes with this change. Suites green: 7/7 ThreadX
configurations, 5/5 SMP.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-16 14:01:07 -04:00
Frédéric Desbiens b93d1ee92f Fixed the simulator ports and thread create paths so they compile when TX_MISRA_ENABLE is defined (#742)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
r52_fvp / r52 (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
* Fixed the simulator ports so they compile when TX_MISRA_ENABLE is defined

_tx_thread_stack_build() in the four simulator ports converts the fake stack
pointer through TX_POINTER_TO_ALIGN_TYPE_CONVERT and
TX_ALIGN_TYPE_TO_POINTER_CONVERT. tx_api.h defines both macros only in the
non-MISRA branch of its #ifdef TX_MISRA_ENABLE, so with that macro defined the two
names are undeclared and none of the four files compiles.

Reproduced with:

    gcc -m32 -c -DTX_MISRA_ENABLE -I common/inc -I ports/linux/gnu/inc \
        ports/linux/gnu/src/tx_thread_stack_build.c -o /dev/null

which reports both names as implicit declarations and then an int to pointer
assignment. The same command without the define compiles cleanly.

No build configuration under test/tx/cmake or test/smp/cmake defines
TX_MISRA_ENABLE, so CI never compiles these files in that mode. It surfaced on a
branch that carries such a configuration.

The conversions are now written inline, which is what the non-MISRA macros expand
to and what the surrounding port code already does, including the line this
replaced.

The alternative would be the idiom common/src/tx_thread_create.c uses for the same
conversion: an explicit #ifdef selecting the ULONG pair under MISRA. That is not
equivalent here. On __x86_64__ this port defines ULONG as unsigned int and
ALIGN_TYPE as unsigned long long, so a pointer round-tripped through the ULONG pair
loses its top 32 bits. Writing the conversion inline keeps one form that is correct
in both modes and on both widths.

Verified by compiling the Linux port with and without TX_MISRA_ENABLE, both clean.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Fixed the same MISRA build break in the SMP and module manager thread create

_tx_thread_create() in common_smp and _txm_module_manager_thread_create() convert the
thread's stack start through TX_POINTER_TO_ALIGN_TYPE_CONVERT and
TX_ALIGN_TYPE_TO_POINTER_CONVERT with no conditional at all. tx_api.h defines both
macros only in the non-MISRA branch, so with TX_MISRA_ENABLE and
TX_ENABLE_STACK_CHECKING both defined neither file compiles. It is the same defect as
the simulator ports in the previous commit, in two more files.

Reproduced with:

    gcc -c -DTX_MISRA_ENABLE -DTX_ENABLE_STACK_CHECKING \
        -I common_smp/inc -I ports_smp/linux/gnu/inc \
        common_smp/src/tx_thread_create.c -o /dev/null

which reports both names as implicit declarations. Both files now compile with and
without TX_MISRA_ENABLE.

The conversions are written inline for the same reason as the ports: the #ifdef idiom
that common/src/tx_thread_create.c uses selects the ULONG pair under MISRA, which
truncates a pointer wherever ALIGN_TYPE is wider than ULONG. That is reported
separately.

Verified: the SMP regression suite passes 117 of 117, and the ThreadX suite still
builds.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-15 23:19:32 -04:00
Frédéric Desbiens 9e4c57138d Normalized the AI disclosure comment to one fixed line per file (#740)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
r52_fvp / r52 (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The per-edit disclosure named the product and model, so every agent and every
model version appended another line. 74 files carried two to four of them, and
the same five products had accumulated 13 spellings -- Copilot against GitHub
Copilot, Claude Sonnet 4.6 against claude-sonnet-4.6, four spellings of Codex.
Twenty assembly lines carried a doubled comment marker, `; //` or `@ //`.

Every file now carries exactly one line, fixed text naming no product:

    Portions of this file were generated with AI assistance.

It is written with the comment character that file already uses, so the `;`
and `@` assembly files keep theirs and the doubled markers are gone. Precise
attribution stays on the commit, where the Assisted-by trailer is per-change,
dated and attached to the diff it describes. A header line cannot hold that
record honestly, because the code it names gets rewritten and the line stays.
A file-level flag answers whether; the history answers who.

Comment-only. 455 files, 455 insertions and 574 deletions: every removed line
was a disclosure line, every added line is the fixed text, and no file is left
with zero or with more than one. `scripts/check_ports.sh` passes, including the
reproducibility check that would catch a ports_arch master and its generated
copies drifting apart. Recompiled against dev, every file that builds without a
vendor toolchain gives a byte-identical object: 19 of 19 C files under common,
100 of 100 GNU assembly files, and all 16 assemblable files whose comment
marker changed. The 10 remaining marker changes are ac5 and IAR sources where
`;` already started the comment and only the redundant `//` was removed.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-15 17:17:26 -04:00
Frédéric Desbiens f5e59a1dd3 Completed Windows simulator support and regression coverage (#736)
Completed and stabilized the Win32 and Win64 MSVC simulator ports.

- Replaced high-latency host synchronization with bounded scheduler handoffs and
  critical sections.
- Added a high-resolution timer, tick batching, idle fast-forward, shutdown
  coordination and Win64 extension-pointer support.
- Brought the Windows regression tooling and the new thread-transition tests up
  on CMake, Ninja and the Visual Studio Build Tools.
- Extended the SMP teardown diagnostics and corrected 64-bit trace-test handling.

This supersedes the historical `win64`, `win32-perf`, `windows-sim-ports` and
`windows-sim-ports-completion` branches; no unmerged change from them is missing
here. The original Win64 port landed in #529.

1,610 of 1,610 tests pass, across five configurations each: Win32 515, Win64 515,
Win64 SMP 580. No external dependency was added, and the existing MSVC warnings
in the trace configuration are unchanged. Hardware validation does not apply to
host simulator ports.

Assisted-by: Codex (gpt-5.6-sol) <codex@openai.com>
2026-09-15 16:47:08 -04:00
Frédéric Desbiens 9b2979e6b0 Replaced the GNU-only dsb/isb 0xF operands with the UAL sy form (#729)
Fixes #551

`_tx_thread_system_return_inline()` in the Cortex-M `tx_port.h` headers spells
its barriers `dsb 0xF` and `isb 0xF`. A bare hexadecimal operand is a GNU
assembler extension, and IAR rejects it with `operand syntax error`, so the
header cannot be included at all. The block is guarded for GCC, armclang and IAR
together, so every IAR user of an affected port hits it -- four independent
reports on Cortex-M33 and M7 with EWARM 9.50 and 9.70.

Both operands become `sy`, the Arm UAL name for exactly what `0xF` encodes. The
generated instruction is unchanged. Applied to the two `ports_arch` masters and
all 32 copies under `ports`, covering M0, M23, M3, M33, M4, M52, M55, M7 and M85
across ac5, ac6, gnu, iar and keil, plus the `scripts/check_ports.sh` probes that
matched the old spelling.

`check_ports.sh` passes, the copy scripts still reproduce every generated port
byte for byte, and `arm-none-eabi-gcc -O2` compiles a caller for every patched
header, emitting `dsb sy` and `isb sy`. Three headers that need toolchain
intrinsics GCC does not ship fail identically on `dev`.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-15 16:24:45 -04:00
Frédéric Desbiens 5d235a534c Fixed the garbage _tx_initialize_unused_memory in the GNU Cortex-A ports (#726)
Fixes #435

`LDR x1, =__top_of_ram` already loads the top of RAM, so the `LDR x1, [x1]`
that followed read whatever sat at that address and left
`_tx_initialize_unused_memory` holding garbage.

Dropped that instruction from the 13 non-SMP GNU Cortex-A ports, the ARMv8-A
source they are generated from, and the two Cortex-A35 module examples. The SMP
GNU ports already had the correct form, and the Arm Compiler ports are
unaffected -- their symbol really does need the dereference.

`scripts/check_ports.sh` passes and the `gnu` CI job build-verifies the AArch64
ports. Not run on hardware.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-15 16:24:09 -04:00
Frédéric Desbiens dde43b8ab2 Replaced the stale system stack switch pseudo-code in the ARMv7-A ports with comments that describe what the code actually does (#735)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The context save, vectored context save and system return routines in the
ARMv7-A ports carried pseudo-code comments claiming that they saved the
thread stack pointer and then switched to _tx_thread_system_stack_ptr.
Neither of those things happens, and none of these ports references that
variable outside an unused IMPORT in their example builds.

On ARMv7-A each processor mode has its own banked stack pointer. The IRQ
handler branches to _tx_thread_context_save while still in IRQ mode, so the
core's banked IRQ stack already serves as the system stack, and the thread
stack pointer is stored in the control block by _tx_thread_context_restore,
and only when the interrupt results in preemption. The scheduler runs on the
banked SVC mode stack that the startup code sets up. There is nothing for a
software stack switch to do.

The comments were therefore misleading rather than merely redundant, and had
led at least one user to try to restore the code they described. They are now
replaced by a description of the actual mechanism.

The AArch64 SMP ports keep their comments unchanged, because ARMv8-A does not
bank a stack pointer per processor mode and those ports do reload
_tx_thread_system_stack_ptr[core] explicitly.

This is a comment-only change. Every changed line is a comment, and all
twenty-five GNU variants still assemble cleanly for their target core.

The fourteen files under ports/cortex_a{5,7,8,9,12,15,17} were regenerated
from ports_arch/ARMv7-A/threadx/common/src/tx_thread_system_return.S with
ports_arch/ARMv7-A/update.sh. The ARMv7-A SMP ports have no generator, so
those files were edited directly.

Fixes #734

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-11 13:27:32 -04:00
Frédéric Desbiens 7f6e496233 Added the missing VFP enable field to the Cortex-R5/AC5 thread control block, which VFP builds were writing over the FileX pointer (#715)
The Cortex-R5/AC5 port defines TX_THREAD_EXTENSION_2 as empty while its
assembly reads and writes the per-thread VFP enable flag at [thread, #144].
With TX_ENABLE_VFP_SUPPORT that offset lands on tx_thread_filex_ptr, so
tx_thread_vfp_enable corrupted the FileX pointer and lazy save/restore
tested an unrelated value.

Defined TX_THREAD_EXTENSION_2 as ULONG tx_thread_vfp_enable, matching the
Armv7-A ports and the Cortex-R4/R5 GNU and AC6 ports fixed earlier, and
added tx_port_offset_check.c so a future layout change breaks the build
instead of silently retargeting the accesses. The offset was measured at
144 for this port and asserted.

Fixes #382

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-09 17:17:37 -04:00
Frédéric Desbiens 383dd311cd Removed the vector table offset register and system stack pointer setup from the Cortex-M low-level initialization (#714)
_tx_initialize_low_level wrote VTOR and derived _tx_thread_system_stack_ptr from
the reset vector on every Cortex-M port. Programming VTOR is the job of the
low-level startup code that runs before the kernel is entered, and doing it again
inside ThreadX silently overrode a vector table that the application, a
bootloader or a firmware update had already installed. It also forced every
application to export a vector table symbol that the kernel itself never needed.

_tx_thread_system_stack_ptr is never read by any Cortex-M port, so the value
copied out of the reset vector served no purpose.

Both blocks are removed from all Cortex-M ports, together with the now-unused
symbol declarations. The GreenHills files keep the NVIC base address load that
the removed block used to leave in r0 for the SysTick setup that follows.

Applications that relied on ThreadX programming VTOR must now set it in their
startup code. The bundled examples link their vector table at address zero and
run on the reset default.

Fixes #370

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-09 16:28:59 -04:00
850a172bac Added lazy FPU stacking and QEMU functional tests for RV64 GNU port (#549)
* add lazy FPU stacking to context save/restore

Save mstatus/sstatus to stack slot 29 and skip floating-point register
save/restore when FS is Off (bits 14:13). This avoids unnecessary FP
context work for threads that do not use the FPU.

- context_save: check FS in nested and first-level interrupt paths
- context_restore: gate FP restore on nested, no-preempt, and preempt paths
- use sstatus when TX_RISCV_SMODE is defined, otherwise mstatus

* add QEMU virt CMake build and automated test runner
Wire the QEMU virt demo into the CMake build system and add a
Python/GDB functional test runner, mirroring the risc-v32/gnu port.
- Add qemu_virt/CMakeLists.txt to build kernel.elf and register the
  check-functional-riscv64 target (requires Python3; skipped if absent)
- Link kernel.elf with --whole-archive so all ThreadX symbols resolve
- Pin _start at 0x80000000 via .text.boot in entry.s and
  KEEP(*(.text.boot)) in link.lds
- Extend demo_threadx.c with fpu_test_val and shorten thread_0 sleep
  for GDB-driven FPU, timer, and preemption checks
- Add test/azrtos_test_tx_gnu_riscv64_qemu.py; verified passing on
  QEMU virt (FPU, timer interrupt, preemption)

* Clean up RV64 PR scope and remove QEMU test integration leftovers

Revert accidental RV64 qemu_virt test/CMake integration changes and keep this branch
focused on lazy FPU context handling only. Also remove unintended TX_RISCV_SMODE-based
mstatus/sstatus save path and align comments/logic to mstatus-only behavior.

* Initialize mstatus.FS in RV64 stack build so new threads start with clean FP state

Slot 29 was left uninitialized while context restore reads it as an FP-live
hint; garbage FS bits could make a new thread inherit the previous thread's
floating-point registers.

* Add RV64 regression test for the FP state of a newly created thread

The test dirties every floating point register, then creates a thread and
checks that the stack builder wrote the mstatus slot and that the new
thread starts with all floating point registers zeroed. It is registered
for RV64 only, since the RV32 stack builder still leaves the slot unwritten.

* Completed the RISC-V64 lazy FPU so the restore side matches the save side

The lazy FPU save in this branch skips the floating-point stores when
mstatus.FS is Off, and records the mstatus it judged that on in frame slot
29. Merged onto current dev, only the save side had that treatment: both
restore paths and the scheduler's interrupt-frame path still reloaded the
FP registers unconditionally, from slots the save had deliberately left
alone. A thread that never touched the FP unit would have had whatever the
frame happened to contain loaded into its registers, and FS driven to
Dirty on the way out.

The guard is added at the three places that consume an interrupt frame:
both paths in _tx_thread_context_restore, and _tx_thread_schedule_loop.
Each reads slot 29 and skips the FP block when FS was Off, which is the
same shape the risc-v32 port already uses.

The solicited path is deliberately left alone. _tx_thread_system_return
saves the callee-saved FP registers unconditionally, so restoring them
unconditionally is consistent; making that pair lazy as well is a separate
change, and risc-v32 is the model for it.

Verified with QEMU on all five configurations:

  risc-v64   96 of 96 passing, five configurations, nothing unlinkable
  risc-v32   95 of 95 passing, five configurations, unchanged
  functional check-functional-riscv64 passes every check

The ninety-sixth test is the one this branch adds. It is load bearing:
seeding stack build with FS = Off instead of Initial makes it fail, and
restoring the seed makes it pass, so it guards the behaviour the rest of
this branch is about rather than passing regardless.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

---------

Co-authored-by: r <r@r>
Co-authored-by: Frédéric Desbiens <frederic.desbiens@eclipse-foundation.org>
2026-09-09 16:25:28 -04:00
Frédéric Desbiens 6ce8d5cc76 Marked every published ThreadX include directory as SYSTEM so applications no longer get warnings from ThreadX headers (#713)
* Marked every published ThreadX include directory as SYSTEM so applications no longer get warnings from ThreadX headers

Commit 8c3c08f added the SYSTEM keyword to the target_include_directories
call in common/CMakeLists.txt, but the same call in every port, in the
SMP common directory's consumers, in the POSIX and FreeRTOS compatibility
layers, and for the generated tx_user.h directory was left unchanged.
A consumer building with a strict warning set therefore still saw
diagnostics coming from tx_port.h and from the compatibility layer
headers, which is exactly what the original change set out to avoid.

Added the SYSTEM keyword to all of those calls so CMake emits -isystem
rather than -I for every directory holding a ThreadX public header. The
ARMv7-M, ARMv8-M, ARMv7-A and ARMv8-A architecture sources under
ports_arch were updated alongside the ports they generate, keeping the
two in step.

Directories that are PRIVATE to an example or test build were left as
they are, since nothing is published from them.

Fixes #290

Assisted-by: Copilot (Opus 5) <noreply@github.com>

* Stopped apt-get update being a gate it was never meant to be

apt-get update fails if any configured repository serves a bad index,
including ones this project never reads. The GitHub runner image carries
Google's and Microsoft's apt repositories, and a Hash Sum mismatch from
Google's, their CDN caught mid-publish with the index and the Release
file eight hours apart, failed all three attempts and turned a run red
over a browser nobody was installing.

Made a failed update warn and carry on, leaving apt-get install as the
gate. Nothing is weakened by that: the install still exits on a package
it cannot find, so an unreachable archive still stops the script, one
step later and naming the package it could not get, which is a better
diagnostic than a hash mismatch in a repository nobody asked for.

Disabling third-party sources before updating would keep the update
strict, but this script also runs on a contributor's own machine, and
rewriting someone's apt configuration to suit CI would be worse than
tolerating a stale index for an archive we do not read.

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-09 14:27:02 -04:00
4d90a0c21c Stopped the win32 and win64 ports from enabling performance metrics and event trace (#676)
* win32: do not always enable trace or performance metrics in tx_port.h

if these are required they can be enabled in tx_user.h

* win64: do not always enable performance metrics in tx_port.h

if required they can be enabled in tx_user.h

* Removed the disabled blocks rather than commenting them out

The win32 and win64 ports were the only two that turned performance
metrics on for the application, and win32 the only one that turned event
trace on. Leaving those to tx_user.h is right: the symbols extend the
control blocks, so a port that sets them behind the application's back
changes structures the application also sees.

The blocks were disabled with #if 0 rather than deleted. That is the form
MISRA C:2012 Directive 4.4 is about -- sections of code should not be
commented out -- and it leaves two copies of a list that now has no
reader. They are removed, and a short note in their place says where the
symbols belong and why the port does not set them.

No behaviour change beyond what this pull request already made. Checked
that the preprocessor nesting in both headers is still balanced.

Worth recording for whoever looks next: with this in, no port defines
either symbol. The linux port carries the same list commented out, which
reads at a glance like a third case but is not one.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

---------

Co-authored-by: r <r@r>
Co-authored-by: Frédéric Desbiens <frederic.desbiens@eclipse-foundation.org>
2026-09-09 13:39:13 -04:00
Akif EjazandFrédéric Desbiens 40db27e843 riscv32: spec compliance and regression test fix (#691)
* riscv32: spec compliance and regression test fix

Signed-off-by: Akif Ejaz <akifejaz40@gmail.com>

* Derived the RISC-V32 frame sizes from the port contract in one place

tx_port.h published TX_RISCV_TRAP_FRAME_SIZE for the GNU BSP assembly, but
nothing in the port consumed it. Six .S files each rebuilt the same numbers
from their own #if, so the interrupt frame size was written out in seven
places and the solicited frame size in three.

That is the shape that produced the RISC-V64 fault fixed in #708, where the
port moved to a padded frame and one copy of the constant did not. The
sources now include tx_port.h and take both sizes from it, and no literal
frame size remains in the port. TX_RISCV_SOL_FRAME_SIZE joins the contract,
since the solicited frame was never published at all.

The emitted code is unchanged: 400 and 176 bytes for ILP32D, 128 for
soft-float, confirmed by disassembly before and after.

Two further corrections:

_tx_initialize_low_level carried .global immediately followed by .weak, so
the symbol stayed weak and the .global did nothing. Weak is what the port
wants, because the example and regression BSPs both provide their own
definition, so the stray .global is removed rather than the .weak. Verified
with nm that the symbol is still W.

The QEMU runner seeded fpu_verified from skip_fpu, so a soft-float run
satisfied the FPU gate whether or not the script ever reported the skip. It
now starts false and is set only when the skip marker is present, so a run
that dies before reaching that point fails instead of passing.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

---------

Signed-off-by: Akif Ejaz <akifejaz40@gmail.com>
Co-authored-by: Frédéric Desbiens <frederic.desbiens@eclipse-foundation.org>
2026-09-09 11:40:05 -04:00
Frédéric Desbiens 3fc28d8979 Allowed a BASEPRI-masked interrupt to wake the Cortex-M idle loop from WFI (#711)
The idle loop in _tx_thread_schedule masks interrupts before re-reading
_tx_thread_execute_ptr, so that an interrupt cannot make a thread ready
between that read and the WFI and then be lost. With PRIMASK that is safe:
the architecture excludes PRIMASK from the WFI wake-up condition, so the
interrupt still wakes the processor and is taken as soon as the loop
re-enables interrupts.

BASEPRI is not excluded. When TX_PORT_USE_BASEPRI is defined, the mask
written at the top of the loop is still in place at the WFI, so any
interrupt at or below TX_PORT_BASEPRI fails to wake the processor at all.
On a system where every interrupt is managed by ThreadX, that is every
interrupt, and the processor stays in WFI or in the low power mode entered
by tx_low_power_enter until something outside the mask happens.

Set PRIMASK and clear BASEPRI for the duration of the WFI, then restore
BASEPRI and clear PRIMASK. Interrupts remain masked across the whole
window, so the original race is still closed, but the wake-up condition is
now evaluated with BASEPRI clear and any enabled interrupt can end the wait.
A pending interrupt is taken at the CPSIE i, or at the existing unmask
further down the loop if it falls below TX_PORT_BASEPRI.

The change is made in the ARMv7-M and ARMv8-M sources under ports_arch,
including the module manager, and propagated to the generated ports with
the copy scripts. The ghs and keil ports do not implement
TX_PORT_USE_BASEPRI, and Cortex-M0/M0+/M23 have no BASEPRI register, so
those ports are unaffected.

Verified by assembling every patched GNU port for Cortex-M3, M4, M7, M33,
M55 and M85, with and without TX_PORT_USE_BASEPRI, TX_LOW_POWER and
TX_ENABLE_EXECUTION_CHANGE_NOTIFY, and by checking the disassembly of the
idle loop.

Fixes #279

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-09 11:29:14 -04:00
Frédéric Desbiens 146d57b235 Fixed the RISC-V64 trap frame size mismatch in the regression test BSP (#708)
The RISC-V64 port moved its interrupt frame to 528 bytes (65 slots plus 8
bytes of padding, so sp stays 16-byte aligned at a call) and published the
size as TX_RISCV_TRAP_FRAME_SIZE. The port sources were converted to use
it, but the shared regression test BSP was not: its trap_entry still
allocated a hardcoded 65 * REGBYTES, or 520 bytes.

Every interrupt therefore unwound 8 bytes more than it allocated:

    trap_entry:                   addi  sp,sp,-520
    _tx_thread_context_restore:   addi  sp,sp,528

On the RISC-V64 regression suite that left 24 of 95 tests failing in the
default configuration, typically as an illegal instruction once execution
reached a corrupted frame. The example BSP under the port directory was
converted with the port and was unaffected, which is why the functional
QEMU test kept passing.

The test BSP now takes both frame sizes from the port it is linked
against, so the two cannot drift apart again. A port that publishes no
contract keeps the historical layout, so the RISC-V32 side is unchanged
until its own port publishes one.

TX_RISCV_TRAP_CALL_FRAME_SIZE is restored to the RISC-V64 tx_port.h. It
was removed as unused when the frame sizes were introduced, but it is
part of the same contract: it is the space a trap entry reserves around a
call into C, and the psABI requires 16 bytes there rather than one
register slot.

Verified on QEMU with every linkable test built, comparing against the
commit before the port change:

    before the port change   2 failures out of 95 (both unlinkable)
    current dev              24 failures out of 95
    with this change          2 failures out of 95 (both unlinkable)

The two remaining failures predate all of this: newlib pulls _impure_ptr
out of R_RISCV_HI20 range for time(), so those two binaries do not link.
RISC-V32 is unchanged at 2 failures across all five configurations.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-09 10:56:57 -04:00
Frédéric Desbiens d3fb72b6dd Aligned the simulator ports' fake stack pointer so ThreadX no longer performs misaligned ULONG accesses (#705)
The Linux, Win32 and Win64 simulation ports build a fake initial stack
pointer by subtracting a fixed 8 bytes from tx_thread_stack_end. That
field addresses the last byte of the thread's stack area, so it is one
less than an aligned address and the resulting pointer is misaligned by
construction, no matter how well aligned the stack the application
supplied was.

Two ULONG accesses then use that pointer. _tx_thread_stack_build() itself
clears the word below it, and _tx_thread_create() copies it into
tx_thread_stack_highest_ptr, which TX_THREAD_STACK_CHECK dereferences on
every suspend and resume when TX_ENABLE_STACK_CHECKING is defined. Both
are undefined behaviour. They happen to work on x86 but are reported by
GCC's undefined behaviour sanitizer, and would fault on a host that
requires natural alignment.

The fake stack pointer is now rounded down to a ULONG boundary, which
leaves it inside the stack area and makes both accesses aligned.

Verified by building the Linux port with -fsanitize=undefined and
running the demo: the two reported diagnostics are produced before the
change and neither appears after it. The tx and smp regression suites
pass, 98 and 114 tests respectively.

Fixes #218

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-09 09:34:33 -04:00
Akif EjazandFrédéric Desbiens 164f211a01 riscv64: spec compliance and regression test fix (#698)
* spec compliance

Signed-off-by: Akif Ejaz <akifejaz40@gmail.com>

* revert the demo changes

Signed-off-by: Akif Ejaz <akifejaz40@gmail.com>

* Restored the FPU demo hook so the RISC-V64 functional test can pass

The new check-functional-riscv64 target verifies FPU context switching by
watching fpu_test_val advance by 1.1f on each pass through
thread_6_and_7_entry. The GDB script deliberately treats a missing symbol
as a failure rather than silently skipping the check, but the demo no
longer defined it, so the target failed on every run:

    FPU_VERIFIED_FAIL_NO_SYMBOL

The definition and the increment are restored, matching what the risc-v32
demo already carries. The functional target now passes end to end.

Three small corrections are folded in:

- tx_port.h carried a comment stating that the ISA string must include
  Zicsr, but nothing enforced it, so an rv64imac build failed with a wall
  of assembler "unrecognized opcode" errors. It now stops at one clear
  diagnostic.
- Removed TX_RISCV_TRAP_CALL_FRAME_SIZE, which nothing referenced.
- The example .gitignore listed qemu-riscv32.log, but the runner writes
  qemu-riscv64.log, so the generated log showed up as an untracked file.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

---------

Signed-off-by: Akif Ejaz <akifejaz40@gmail.com>
Co-authored-by: Frédéric Desbiens <frederic.desbiens@eclipse-foundation.org>
2026-09-09 09:26:11 -04:00
Frédéric Desbiens 945f5f5caa Moved the ARC ISR enter callout onto the system stack (#700)
_tx_thread_context_save() calls _tx_execution_isr_enter when
TX_ENABLE_EXECUTION_CHANGE_NOTIFY is defined. On the path where an interrupt
preempted a running thread, that call was made before the switch to the system
stack, so the 32-byte call frame and the whole stack footprint of the callout
were taken from the interrupted thread's stack, on top of the 160-byte interrupt
frame the port had already allocated there.

The callout is supplied by the application, so its stack usage is not bounded by
ThreadX, and it is charged to every thread that happens to be running when an
interrupt arrives.

The two other callout sites in the same routine, the nested save and the idle
system save, already run on the system stack, as does the _tx_execution_isr_exit
call in _tx_thread_context_restore. _tx_thread_schedule was reordered in 6.1.9
so that _tx_execution_thread_enter runs on the system stack rather than the
thread stack; the same reorder was never applied to the context save.

The switch to the system stack now happens before the callout in the ARCv2_EM,
ARC_HS and SMP ARC_HS ports. _tx_thread_context_fast_save is unchanged because
the fast interrupt path never switches stacks by design.

Fixes #149

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-08 15:42:39 -04:00
Frédéric Desbiens d7789f0b12 Fixed the clobbered return address in the RISC-V context save (#696)
_tx_thread_context_save() returns to its caller with ret, which uses the
return address held in ra. When TX_ENABLE_EXECUTION_CHANGE_NOTIFY was
defined, the call to _tx_execution_isr_enter overwrote ra with the address
of the instruction following the call, so the subsequent ret returned into
_tx_thread_context_save itself instead of the interrupt service routine.

The return address is now saved on the stack around the call and recovered
afterwards, which is the same idiom already used by the Arm ports. The fix
covers all three affected paths (nested save, thread save and idle system
save) in the risc-v32 GNU, risc-v32 IAR and risc-v64 GNU ports.

Fixes #348

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-03 14:17:49 -04:00
Frédéric Desbiens 508af549da Fixed the incorrect loop bound constant in the IAR file lock support (#695)
The IAR multithreaded library support code allocates its file lock
mutexes from an array of _MAX_FLOCK entries, but the wrap-around check
and the exhaustion check in __iar_file_Mtxinit() both compared against
_MAX_LOCK, the bound of the unrelated system lock mutex array.

When _MAX_FLOCK is greater than _MAX_LOCK, the free mutex index wrapped
early and the exhaustion check reported failure while free entries
remained, so *m was set to TX_NULL and the application faulted the first
time a file lock was taken.  When _MAX_FLOCK is smaller than _MAX_LOCK,
the free mutex index was allowed to run past the end of
__tx_iar_file_lock_mutexes and the exhaustion check could never fire.

Corrected all four comparisons in each of the 27 copies of tx_iar.c.

Fixes #444

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-03 13:50:31 -04:00
Frédéric Desbiens c469a7756b Fixed the missing immediate prefix on MOV in Cortex-M schedulers (#693)
The BASEPRI-masking path in tx_thread_schedule wrote "MOV r0, 0" rather
than "MOV r0, #0". UAL requires the "#" prefix on an immediate operand.
GNU as and the LLVM-based assemblers accept the unprefixed form and emit
the intended encoding, but stricter assemblers reject it outright, so the
affected ports could not be built with those toolchains.

The GNU and AC6 sources had already been corrected; this brings the IAR
and AC5 sources into line. Verified that both spellings assemble to the
same Thumb-2 encoding (f04f 0000), so this is a source-correctness fix
with no change in generated code or runtime behaviour.

Covers 31 occurrences across the Cortex-M3, M4, M7, M33, M52, M55 and M85
ports, their module manager counterparts, and the shared ARMv7-M and
ARMv8-M architecture sources.

Fixes #461

Assisted-by: Copilot (Opus 5) <noreply@github.com>
2026-09-03 11:44:27 -04:00
Frédéric Desbiens 13c8c768c7 Added the boot-at-EL1 option to the S32Z280 entry path (#690)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The Armv8-R AEM FVP entry path has carried TX_R52_BOOT_AT_EL1 since it
was written, for the case its own comment describes: "an earlier boot
stage or a vendor EL2 monitor has already dropped privilege to EL1".
This board's entry path did not, so a kernel could not be built as a
guest on it at all -- and it is the board where that matters most,
because it is the one with silicon behind it.

The bracket is the whole change.  Everything from the Thumb reset
trampoline to the ERET goes inside #ifndef TX_R52_BOOT_AT_EL1, and the
#else supplies a one-instruction A32 _start that branches to el1_entry.

A32 AND NOT T32, which is the one real difference from the standalone
entry.  The core resets in Thumb state here because the RTU boot
instruction NXP plants is a T32 branch, but a guest is not reached by
reset: it is reached by the monitor's ERET, and the monitor chooses the
state through SPSR.T.  Get the two out of agreement and the guest dies
on its first instruction with an undefined-instruction exception, which
looks exactly like a bad entry address and sends the reader to the
loader instead of to the ERET.

WHAT THE MONITOR INHERITS is enumerated at the #ifndef, next to the code
it replaces rather than in a document, because that is where somebody
adding a third board will be looking.  This board's EL2 block is
considerably larger than the model's, and each item on the list is
something a guest at EL1 provably cannot do rather than something it
merely does not: CNTFRQ is writable only at the highest implemented
exception level and reads zero out of reset; HCPTR.TCP10/TCP11 reset
set, trapping every EL1 floating-point access; HSCTLR.TE is an EL2
register (SCTLR.TE is EL1's, and el1_entry still clears it below);
ICC_HSRE.SRE makes every other ICC_* and ICH_* register exist at all;
the low-latency peripheral port enables reset to zero and an EL1 write
to that register traps to EL2; and the TCM enables are per-core with
ENABLEEL2 SILENTLY IGNORED from EL1 -- measured on both BTCM and CTCM,
the base took and bit 0 took while bit 1 stayed clear.

CNTHCTL.PL1PCTEN and PL1PCEN are the deliberate omission from that list,
and the note says why.  This path opens both, because a standalone
kernel owns the physical timer.  A monitor that TIME-partitions its
guests must not: a partition's physical time keeps running while it is
descheduled, so a guest reading it can observe that it was not running.
That is the monitor's decision rather than this file's, which is why the
list says what a guest cannot do rather than what a monitor should.

Verified both ways.  The three standalone images build and link
unchanged, and a kernel built with the option boots at EL1 on a
S32Z280-594EVB under an EL2 monitor, runs two threads through a queue
and a semaphore, and reports back -- with no other change to the kernel
or to its port.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-02 19:08:24 -04:00
Frédéric Desbiens 57390fc0fe Refused the Cortex-R52 VFP option without a hard float ABI (#686)
TX_R52_ENABLE_VFP with the default soft float ABI cannot build.  The option
defines TX_ENABLE_VFP_SUPPORT, which enables the VMRS, VSTMDB and VLDMIA
blocks in the port assembly, and -mfloat-abi=soft leaves the assembler with
no FPU to accept them.  The configuration fails with eight errors of the form

  tx_thread_system_return.S:118: Error: selected processor does not support
  `vmrs r4,FPSCR' in ARM mode

none of which mentions the float ABI, so a user has to reason from VMRS back
to the option that enabled it.

The guard that exists said otherwise.  It warned that "the compiler will not
emit floating-point instructions, so the VFP context path will never be
exercised", which describes a build that succeeds and is merely pointless --
and then let configure finish, so the warning scrolled past well before the
assembler errors appeared.

It is now a FATAL_ERROR that names the fix, which is what the same file
already does six lines above for TX_R52_ENABLE_FIQ_NESTING without
TX_R52_ENABLE_FIQ.  That combination is rejected for being "meaningless",
while this one, which cannot assemble at all, was only warned about.  The
severities were the wrong way round.  The option's definition also moves
below the check, so the block reads like the FIQ nesting one.

The ABI is not promoted to hard automatically.  TX_R52_FLOAT_ABI is a cache
variable the user may have set deliberately, and silently overriding an
explicit choice is worse than refusing a combination that cannot work.

readme_threadx.txt carried the same claim, and its option list marked the FIQ
nesting dependency inline but not this one.  Both corrected.

No regression test.  Nothing in the tree asserts a configure-time failure --
there is no harness for it, and the sibling FIQ nesting guard has none either
-- so a test for this would have to introduce that mechanism for one case.
Verified by hand in both directions instead: the soft-ABI combination now
stops at configure with the message above, and the hard-ABI feature build
(VFP, FIQ, IRQ nesting, FIQ nesting) builds its eight images clean and passes
ctest 8/8 on the Armv8-R AEM FVP.  scripts/check_gcc.sh passes unchanged; it
configures the Cortex-R52 CMake stage without TX_R52_ENABLE_VFP, so the new
branch is not on its path.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-01 08:48:20 -04:00
Frédéric Desbiens 5b94bad6a2 Fixed the AArch64 samples, none of which had ever linked with GCC (#673)
Every AArch64 gnu example build failed at the sample link, all 27 of them --
13 under ports/ and 14 under ports_smp/:

  libg.a(libc_a-init.o): in function `__libc_init_array':
      undefined reference to `_init'
      relocation truncated to fit: R_AARCH64_CALL26 against undefined
          symbol `_init'
  libg.a(libc_a-fini.o): in function `__libc_fini_array':
      undefined reference to `_fini'

build_threadx_sample.sh links with -nostartfiles, which is correct for a port
carrying its own reset path, and that drops crti.o and crtn.o along with
everything else. startup.S calls __libc_init_array by design, and newlib's
implementation calls _init, which crti.o is what defines. The AArch32 scripts
are unaffected: they use nosys.specs and never reach __libc_init_array.

The fix links crti.o and crtn.o explicitly, bracketing the object list -- the
first must precede every .init contribution and the second must follow all of
them, so their position is load-bearing rather than stylistic. Both paths come
from the compiler's own -print-file-name, so nothing here hard-codes a
toolchain layout.

The atfe branch sets both to empty, deliberately: picolibc's __libc_init_array
does not call _init, those 27 images link today, and adding crti.o would change
a working link for no reason. That is also why check_clang.sh is green on these
and does not list them as expected to fail -- the LLVM path never reached the
gap, so nothing has ever linked them and failed.

Fixed in ports_arch/ARMv8-A/threadx/ports/gnu/example_build, which is the
single source for both the ports/ and ports_smp/ copies, then regenerated with
update.sh --port-sets tx,tx_smp. The 27 generated copies are in this commit
because ports_arch_check compares them.

Verified: all 27 link with arm-gnu-toolchain 14.3.rel1 aarch64-none-elf, where
0 of 27 did before; _init and _fini disassemble to the expected crti prologue
and crtn epilogue over a ret; check_clang.sh with ATfE 22.1.0 is still green on
all five stages, including the 42 script-driven example builds; check_ports.sh
is green including the reproducibility check.

No regression test: these are link-only example images that no host test
executes. What guards them is check_clang.sh's example stage today, and
check_gcc.sh's, which is the next change and is the reason this was found.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-28 09:26:12 -04:00
Frédéric Desbiens ca62edd27d Swept the stack-heavy measurement across placements, and qualified its result (#638)
#636 reported that a stack in BTCM gave a threefold tighter spread than DRAM0
for stack-heavy work. That measurement used a single code placement, which is
the methodology #631 and #633 exist to correct: the cache benchmark got an
alignment sweep and the interrupt handler got one, and this measurement never
did. It was noticed when #637 added two threads to the same image and the figure
moved -- both spreads came out near 6500 and the minima rose 15%.

The recursive body is now generated at four placements and all four are
measured, per placement, in one image.

    placement    BTCM min / spread     DRAM0 min / spread
    offset 0      41854 / 6850          42036 / 6880
    offset 16     48670 / 1946          48792 / 1978
    offset 32     42388 / 6914          42752 / 7018
    offset 48     48914 / 1860          49198 / 6786

Reproducible across runs to within a few hundred cycles.

Spread is dominated by code placement rather than by the memory holding the
stack. It ranges from 1860 to 6914 depending on where the body falls in a cache
line, and placement also moves the minimum by 17%, from 41854 to 49214. Against
that, the memory contributes a consistent but small advantage: BTCM's minimum is
lower at all four placements, by 0.4% to 0.9%.

BTCM's spread beats DRAM0's decisively at one placement of the four, offset 48,
at 1860 against 6786. At the other three the two are within 2% of each other.
So the effect #636 reported is real where it occurs and is not a property of the
part: quoting it as one invited the reader to expect it everywhere.

#636's claim should be read as qualified by this. A stack in BTCM buys a small
consistent improvement in the best case and a large improvement in spread at
some code placements and not others. Anyone building a determinism argument on
it needs the placement sweep in the loop, not a single figure.

The pad nops that displace each placement execute on every recursion level
rather than once, so each placement carries a slightly different constant cost,
about 0.6% at the widest. That cancels in the BTCM against DRAM0 comparison,
which is made at the same placement, and does not affect spread within one.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-18 08:51:49 -04:00
Frédéric Desbiens 1ef49338e9 Gave each thread its own MPU window, and made a violation fault (#637)
First step towards a ThreadX module port for this core: establish that PMSAv8-R
regions can be switched per thread on this part, what that costs, and that a
violation actually faults. Those are the questions worth answering before
writing a module manager on top of them.

Two threads each own a 4 KB window at the top of DRAM2. The windows are carved
out of the broad data region in mpu.c, because isolation is only meaningful in
memory no other region already covers -- every other region in that map is a
wide RW window, so a private buffer inside one of them would be reachable by
every thread whatever else was programmed.

Each thread writes its own window, which must succeed, and then reaches for the
other thread's, which must fault. The second half is the part that matters: a
test that only shows a thread reaching its own memory would pass just as well
with no protection at all.

Measured on the S32Z280-594EVB, reproducible across three runs:

    thread 0 window 0x3187E000  own: reachable  other: faulted
    thread 1 window 0x3187F000  own: reachable  other: faulted
    region switch cost: 562 to 604 cycles

The cost is worth noting for the module port to come. A context switch on this
part is about 1400 cycles, so switching one region adds roughly 40% to it, and
most of that is the dsb and isb rather than the register writes. A module switch
programming several regions should therefore batch the barriers once at the end
rather than per region.

Scope, stated plainly. The window is applied by the thread calling
thread_mpu_activate, not by the scheduler. The port's scheduler does call
_tx_execution_thread_enter under TX_ENABLE_EXECUTION_CHANGE_NOTIFY, which would
make it automatic, but that macro is read by port assembly compiled into the
shared threadx library, so enabling it would oblige all nine example targets in
this port to supply the four execution hooks. A ThreadX module port carries its
own copies of the port assembly for exactly that reason, and that is where the
switch belongs. There is no user mode, no syscall boundary and no loader here.

The fault is survivable the same way the boot probes make it survivable:
fault_expected tells the data abort handler to record the violation and resume
after the faulting access. That works in thread context because the handler
returns where it came from rather than to a fixed recovery point.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-17 17:59:12 -04:00
Frédéric Desbiens f535a4ec67 Measured stack-heavy work against the memory holding the stack (#636)
#635 found BTCM worth about 7.4% on a context switch with no determinism
advantage, and said why: a switch saves sixteen registers, roughly one cache
line, so the stack's cache state has almost nothing to contribute. It named the
interesting case as work with a large stack working set, said it had not been
measured, and said it should not be assumed. This measures it, and the answer
inverts the earlier one.

deep_touch recurses 24 frames, writing a frame on the way down and reading it on
the way up, so the working set is the whole descent. The cache is cleaned and
invalidated before each sample, so every descent starts cold. Two threads, one
stack in BTCM and one in DRAM0, no partner threads and no relinquish: the timed
region is entirely within one thread.

Reproducible across runs:

                      min      mean      max     spread
    stack in BTCM   42868     43031    44890       2020
    stack in DRAM0  42982     43291    49842       6860

The mean is the same to within 0.6%. The worst case is 10% lower for BTCM and
the spread is 3.4 times tighter. No sample in either configuration exceeded
twice the minimum, so these maxima are the workload rather than a timer tick --
which is the mistake that produced a false jitter result in #635 and is why the
count of interrupted samples is printed.

Put beside #635 the two measurements say opposite things and both are true.
For a context switch, a small footprint touched every time, BTCM buys throughput
and no determinism. For stack-heavy work, a large footprint touched once, it
buys determinism and almost no throughput.

The reason is that this workload is compute bound at the optimisation level this
BSP builds at: 43000 cycles for 24 frames is dominated by call and loop overhead,
so line fills are a few percent of the total and barely move the mean. What they
do is vary, and that variance is what a bank with no cache in the path removes.

So the determinism argument for TCM holds here, but it is worth 10% of worst case
and a threefold narrowing of spread, not an order of magnitude. Anyone citing
this in a safety argument should cite those numbers and not a larger claim.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-17 17:26:08 -04:00
Frédéric Desbiens 367d91880b Measured context-switch cost against the memory holding the stack (#635)
#634 placed a thread stack in BTCM and deliberately claimed no timing benefit,
because none had been measured. This measures it.

Two pairs of equal-priority threads hand control back and forth with
tx_thread_relinquish. One pair has both stacks in BTCM, the other in DRAM0, and
the measuring thread of each pair times the round trip in PMU cycles. Both pairs
run in one image from one copy of the measuring code, which is what makes the
comparison safe: the alignment trap that invalidated earlier work here bites
when two builds with different layouts are compared, and a code shift moves both
pairs equally.

Reproducible to the cycle across runs:

                        min    mean    max
    BTCM stacks        1370    1379   1402
    DRAM0 stacks       1476    1480   1508

BTCM is about 7.4% faster, or 110 cycles on a round trip of two switches.

Three findings that bound the claim, and the last one deflates it.

The figure holds whether the cache is warm or cold. Cleaning and invalidating
the data cache before every timed switch costs both configurations about 40
cycles and leaves the gap at 7.4%: warm it is 1334 against 1440, cold 1370
against 1476. So the advantage comes from BTCM's zero wait states, not from
avoiding cache misses.

That is because a context switch touches almost no stack -- sixteen registers,
about one cache line -- so the stack's cache state has little to contribute
either way. TCM should matter much more for threads with deep call chains or
large locals, where the stack working set is big enough for cache state to
dominate. That is not measured here and should not be assumed.

There is no determinism benefit visible in this test. Excluding preempted
samples, jitter is 32 cycles for BTCM and 28 to 34 for DRAM0 -- comparable, not
better. A first version of this measurement appeared to show BTCM with 16 times
less jitter, and that was wrong: max was reporting whichever pair a timer tick
had landed on. Across three runs the outlier appeared in the BTCM pair once and
the DRAM0 pair twice. Samples past 2000 cycles are now counted separately and
excluded from min, mean and max alike, and the count is printed so the reader
can see how many there were.

Also tried and discarded: loading the partner thread with a cache walk to create
pressure. The timed round trip includes the partner, so the walk dominated every
sample and put all 256 past the outlier threshold. The per-sample flush replaced
it and sits outside the timestamps.

The demo also starts the PMU cycle counter, which bsp_boot.c does for the probe
image and this image never ran. Without it every reading would have been zero,
which reads as a free context switch rather than as a dead counter.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-17 17:08:29 -04:00
Frédéric Desbiens 62e966f6bf Enabled BTCM, measured it, and put a ThreadX thread stack in it (#634)
* Enabled BTCM and measured what it offers as a data store

BTCM was disabled because of a regression that turned out not to exist; the
claim was retracted in the previous commit. It is enabled now, and this
measures why that is worth doing: 16 KB at zero wait states, where ATCM has
one, and no cache in the path at all.

Enabling it needs three things, all of which existed for ATCM already: the
region register write at EL2, an MPU region, and an ECC preload before any read
(TRM 6.2.2). BTCM accepts 32-bit stores where ATCM needs 64-bit, which
tcm_preload already handles.

The cache benchmark now sweeps three memories rather than one, four loop
alignments each. At the alignments where the loop is not instruction-fetch
bound:

    memory              cold (uncached)   warm (cached)   gain
    DRAM2 half-speed            857,540         651,436   24.0%
    DRAM0 full-speed            797,824         651,297   18.3%
    BTCM  zero wait             694,689         651,369    6.2%

Three things follow.

Warm times are identical across all three memories, within 0.02%. Once the data
cache is working the backing store barely matters, because the working set fits
in it.

Cold times rank as the reference manual predicts: BTCM fastest, then DRAM0,
then DRAM2 at half the core frequency (S32Z2 RM 6.3.6).

BTCM still shows a 6.2% gain when the caches are enabled, and that cannot be
the data cache, because an enabled TCM is Non-cacheable Non-shareable Normal
memory whatever the MPU says. It is the instruction cache on the timing loop.
This probe has always measured both caches together; three memories side by
side is what makes that visible.

The number that matters for placing data in BTCM: uncached BTCM is within 6.6%
of the best cached case, where uncached DRAM0 is 22% off it. Data in BTCM runs
at close to cache-hit speed with no cache to miss, which is the determinism
argument stated as a measurement rather than an assertion.

DRAM2's sweep is unchanged with BTCM enabled, 0 and 0 and 240 and 240, which
independently confirms the retraction.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Put a ThreadX thread stack in BTCM

The code side of TCM was done in #630; this is the data side. One of the demo's
three thread stacks now lives in BTCM and the other two stay in DRAM0, so a run
exercises both paths and a mistake in either shows up.

link.lds gains a BTCM region and a .btcm_bss NOLOAD section, so a stack is an
ordinary C array with a section attribute and the linker checks it fits, rather
than a hardcoded address that silently overflows the bank.

entry.S preloads the whole bank at EL2, and that is not optional. ECC is enabled
on this part, so a TCM location must be written before it can be read (TRM
6.2.2), and a stack is read before the program writes it -- the first context
restore pops what tx_thread_create built into it. The preload has to happen
before any C runs, because the demo images do not run bsp_boot.c, which is where
the ATCM preload lives. 32-bit stores suffice for BTCM where ATCM needs 64-bit.

Why BTCM for a stack: 16 KB at zero wait states where ATCM has one, and never
cached whatever the MPU says about it. Measured in the previous commit, uncached
BTCM comes within 6.6% of the best cached case while uncached DRAM0 is 22% off
it, so stack access runs at close to cache-hit speed without depending on a line
being resident. That is the property a determinism argument needs.

What this commit does not claim: no thread-level timing improvement has been
measured. The case for BTCM here rests on the memory characterisation and on
removing the cache from the path, not on a measured context-switch figure. That
measurement is worth doing and has not been done.

Verified on the S32Z280-594EVB. The demo reports its stack addresses so the
placement is visible rather than implied -- sleeper at 0x30100000 in BTCM,
spinner and judge in DRAM0 -- and passes with 100 ticks, 20 sleeper wakeups and
20 preemptions, so a real thread schedules, preempts and context-switches on a
tightly-coupled-memory stack. The boot image still passes six of six probes.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-17 16:33:55 -04:00
Frédéric Desbiens 35de56853c Retracted the claim that a second TCM bank costs the data cache (#633)
entry.S and readme_s32z280.txt both stated that enabling any second TCM bank
removes all measurable data-cache benefit on this part, and gave five
configurations as evidence: ATCM alone at a 24% cache gain, four combinations
involving a second bank at none. Both concluded the cause was not a particular
bank, not its address and not the ECC preload, but enabling a second bank at
all. Both cited the Cortex-R52 and S32Z2 errata as not covering it and offered
a shared RAM pool between the LLC and the TCMs as an explanation. A defect
report went to NXP on that basis.

It was an artifact of the benchmark. That benchmark was bimodal with respect to
where its timing loop fell inside a 64-byte cache line, reporting either 24% or
nothing at all for identical silicon, and every one of those five
configurations was an edit to entry.S, so every one shifted the code that
followed and moved the loop between modes. Adding two nop instructions
reproduces the "second bank" figure exactly, to the digit. The report to NXP
has been withdrawn.

Re-measured with the alignment sweep added in #631, one bank and two are
indistinguishable:

    loop offset in line      ATCM only      ATCM + CTCM
    0                        gain 0         gain 0
    16                       gain 0         gain 0
    32                       gain 240/1000  gain 240/1000
    48                       gain 240/1000  gain 240/1000

So enabling a second bank costs nothing measurable. The banks stay disabled,
but for the ordinary reason that nothing in this example uses them, and both
texts now say that instead. Enabling one is a single line, with the ECC preload
before any read (TRM 6.2.2) and an MPU region as the only prerequisites, both
already handled for ATCM.

The readme also now states the general point, which outlasts the TCM detail: a
single-figure timing result from this example cannot be compared across builds
unless the timed loop's alignment is controlled, because almost any change
shifts code.

Comments and documentation only; no generated code changes. Verified on the
board regardless, since entry.S was touched: six of six probes pass and the
sweep is unchanged.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-17 13:46:33 -04:00
Frédéric Desbiens 3aaaa6700d Revalidated the ATCM handler result across four placements, and it holds (#632)
The handler comparison in #630 measured one alignment, which is the mistake
that made the cache benchmark in this example report 24% or 0% for identical
silicon. The handler body is now generated at four offsets within a cache
line, all four are measured, and the figures are reported per placement.

The claim survives. Mean cycles for the handler body:

    loop offset      code RAM    ATCM
    0                     523     383
    16                    534     377
    32                    539     375
    48                    521     383

ATCM is faster at every placement, by about 28%, and the two sets of means do
not overlap. Worst case improves as well, 510 against 694.

Two things worth recording beyond the headline.

The handler measurement is only mildly alignment sensitive, 3.5% across
placements in code RAM and 2% in ATCM, quite unlike the cache loop's two
modes. So this comparison was less fragile than the cache one, and #630's
direction was right even though its method was not defensible. The absolute
numbers differ from #630 because the body now sits behind a placement wrapper
that adds a call; the comparison is internally consistent either way.

Both variants also report identical cache sweeps, 0 and 0 and 240 and 240,
which settles the regression this branch's predecessor appeared to show. That
apparent regression was the single-alignment probe moving between its two
modes, not anything about ATCM.

One copy of the logic is kept: the wrappers inline a single always_inline
implementation, so the four placements cannot drift apart.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-17 13:33:43 -04:00
Frédéric Desbiens 1d3f3f8e4c Measured the cache benchmark at four alignments, because one is not enough (#631)
This benchmark was bimodal and reported a single number, which made it worse
than no benchmark. The same workload on the same silicon reports either 24%
cache benefit or none at all, decided only by where the loop falls inside a
64-byte line -- and therefore by any unrelated change that shifts code
ahead of it. Two nop instructions added to entry.S were enough to flip it.

That is not a hypothetical. A run of conclusions drawn from this probe turned
out to be measuring code layout: an interrupt-handler comparison, a claim that
enabling a second TCM bank costs all cache benefit, and a follow-up claim that
what mattered was when the TCM region register was written rather than what it
contained. The last of those was reported to NXP as a defect and has had to be
withdrawn. Enabling CTCM and adding two nops produce identical results, to the
digit, because the only measurable consequence of the enable was the eight
bytes of instructions it added.

Four copies of the loop are now generated at different offsets within a cache
line, all four are measured, and the low and high gains are both reported.
Pinning a single alignment was tried first and is not a fix: it silently picks
one of the two modes -- aligned to 64 the loop sits permanently in the low one.

Measured on the S32Z280-594EVB, reproducing exactly across runs:

    loop offset in line    cold      warm      gain
    0                      890,302   890,035   0%
    16                     890,208   889,976   0%
    32                     857,439   651,390   24.0%
    48                     857,631   651,571   24.0%

The cold pass differs between the modes as well, 890k against 857k, so the
loop is slower even with both caches off. The cold pass is instruction-fetch
bound out of code RAM at half the core frequency (S32Z2 RM 6.3.6), and how the
loop straddles lines decides how much of the data cache's contribution is
visible at all. This probe therefore measures both caches together and always
did; the sweep at least makes the variation visible instead of letting one
arbitrary placement stand in for the part.

C4 now passes if any alignment shows a 10% speedup, and says so explicitly
when the low mode does not, so the sensitivity appears in the log rather than
being discovered later.

Verified: with the sweep in place, adding 0, 8, 12 or 20 bytes of nops to
entry.S leaves the reported low and high gains unchanged. Before it, the same
shifts read 24.0%, 0%, 0% and 0%.

Also adds cache_disable_all, which the sweep needs: cache_enable was one-way,
so a second cold reading in one run was impossible.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-17 13:14:16 -04:00
Frédéric Desbiens b2983c08d3 Ran the interrupt handler from ATCM, and measured what that buys (#630)
The TCM work so far enabled ATCM and left it empty, which buys nothing.
This places code in it and measures the result.

link.lds gains an ATCM region and an .atcm_text section whose run address
is in the bank and whose load address is in CODE. tcm_copy_atcm_text
moves it, using 64-bit stores because ECC is enabled on this part and
ATCM requires them (Cortex-R52 TRM 6.2.2); both ends of the section are
8-byte aligned so there is no narrower tail to leave without check bits.
The copy runs after T4 and T5, which write test patterns to the first and
last words of the bank and would otherwise land on top of the code.

s32z280_atcm.elf is the same image as s32z280_boot.elf with the interrupt
service body placed in ATCM. Both targets exist so the comparison can be
repeated on one board in one session without reconfiguring. The service
routine is split into a timed wrapper that stays in .text and a body that
moves, so the wrapper's own cost appears in both measurements and cancels.

Measured in PMU cycles over 64 samples, caches enabled in both:

                 code RAM    ATCM    change
    min               454     334    -26.4%
    mean              458     340    -25.8%
    max               612     466    -23.9%
    spread            158     132    -16.5%

CNTPCT is not used for this: at 8 MHz it cannot resolve a handler body,
let alone the variation in one.

The level shift is the solid part. ATCM is a quarter faster even though
the caches were on and code RAM had the instruction cache available,
which says the handler does not stay resident between interrupts 10 ms
apart -- so each one pays a cold fetch from code RAM, which runs at half
the core frequency where ATCM runs at full speed with one wait state
(S32Z2 RM 6.3.6).

The determinism claim deserves less weight than the numbers first
suggest. The spread narrows by only 16%, and ATCM's worst case still sits
slightly above code RAM's best case, so the two distributions overlap at
the tails rather than separating. Whatever jitter remains is not
dominated by instruction fetch.

Both images pass six of six boot probes.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-17 09:21:27 -04:00
Frédéric Desbiens ff638dae71 Configured the LINFlexD console twice, because once is not enough at -O2 (#629)
The console driver runs its configuration sequence a single time, and that
works only because this BSP is built without an optimisation flag. Compiled
at -O2 the same sequence leaves the line corrupted: every character
partially wrong, in the pattern the file already describes for a
misconfigured module.

What makes it worth guarding against is that the failure is invisible.
UARTCR reads back exactly the value written. LINIBRR and LINFBRR read back
exactly the values written. linflexd_init returns LINFLEXD_INIT_OK. The
registers are right and the line is wrong, so nothing in the returned
status tells the caller the console cannot be trusted.

Localised by bisection: with every other file at -O2 and this one at -O0
the output is clean, and with only linflexd_init at -O2 it is corrupted,
so the fault is in the configuration sequence rather than in the per-byte
transmit path.

The mechanism is not understood, and this commit does not claim to explain
it. Tested and rejected: a 100x larger bound on the wait for
initialisation mode, a settling delay before the first LINSR read, a
settling delay after leaving initialisation mode, a barrier and read-back
between the two UARTCR writes, and waiting for LINSR to report the exit
from initialisation mode. None of those makes a single pass work at -O2.
A second pass does, at both optimisation levels, which is what this does.

Instrumented with a duplicate of the sequence forced to -O2 and reported
through a console repaired afterwards, which is how the register read-backs
above were obtained.

Verified on the S32Z280-594EVB. The boot image passes six of six probes
with the console status still reporting 0x00000000, and the reproducer
builds clean and prints correctly at both -O0 and -O2.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-17 07:59:23 -04:00
Frédéric Desbiens 0e7db2bf79 Stopped the generated Armv8-M readmes claiming a false origin date (#623)
Every Armv8-M port readme ends with

    09-30-2020  Initial ThreadX 6.1 version for Cortex-M85 using GNU tools.

with the core name substituted in by scripts/copy_armv8_m.sh. The date is the
shared port's, so each core inherits it whatever its own history: Cortex-M85 was
announced in 2022 and its readme claims a 2020 origin, and any core added later
gets the same treatment the moment its name joins the generator's list.

The rest of the history block is accurate, since it records changes to the shared
files. Only the closing line asserts something per-core. Reword it to describe
the Armv8-M port itself, and say where a given core's real starting point is.

Regenerating updates the twelve readmes for cortex_m33, cortex_m52, cortex_m55
and cortex_m85 across the three toolchains.

Cortex-M52 makes the point: it arrived in #519 and its readme immediately claimed
a 2020 origin for a core announced in 2023.

The ARMv7-M templates say "Initial ThreadX version 6.1.7 for Cortex-M", with no
placeholder to substitute, so they make no per-core claim and are left alone.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-16 08:42:53 -04:00
Armchina_JidongMeiandFrédéric Desbiens 090f8edc56 Added Cortex-M52 (Armv8.1-M) port (#519)
* ports: add Cortex-M52 (Armv8.1-M) port support

Add ThreadX port for Cortex-M52, supporting three toolchains:
- GNU (GCC)
- AC6 (Arm Compiler 6)
- IAR

Cortex-M52 is an Armv8.1-M Mainline processor sharing the same
architecture profile as Cortex-M55 and Cortex-M85. The port is
functionally identical to the existing Cortex-M85 port.

* ports: update cortex_m52 port copy script,using it update content

-Add cortex_m52 in threadx/scriprts/copy_armv8m.sh
-Using updated copy_armv8m.sh to generate new cortex_m52 port content

* Regenerated the Cortex-M52 port against dev and wired it into the checks

The port was generated from ports_arch/ARMv8-M as it stood on master, which has
since moved on. Rebasing onto dev and running scripts/copy_armv8_m.sh again
brings the twelve stale files into line, which is the point of generating them:
the core picks up every ARMv8-M fix made since without anyone porting it by hand.

Among what it picks up: "MOV r0, 0" becomes "MOV r0, #0" in the schedule and
system-return paths, the non-canonical immediate form that GNU as tolerates and
LLVM's assembler rejects; and gnu/src/tx_initialize_low_level.S goes away, since
the shared source no longer has it.

Two integration points exist only on dev, so the original change could not have
included them.

cmake/cortex_m52.cmake, so the port can be selected the documented way. Every
other Cortex-M core has one. It uses the hard float ABI, as Cortex-M55 and
Cortex-M85 do.

An entry in scripts/check_clang.sh, likewise with -mfloat-abi=hard. That flag is
not decoration: -mcpu=cortex-m52 implies Helium, and building it soft-float ends
in "multilib configuration error: No library available for MVE with soft-float
ABI" on every file, which reads as a broken port rather than a missing flag.

Verified after regenerating: scripts/copy_armv8_m.sh is a no-op, so the tree
matches its source; all 14 assembly sources and 188 C sources compile for
cortex-m52 with Arm Toolchain for Embedded 22.1.0.

Note for anyone building with GNU tools: arm-none-eabi-gcc 13.2.1 rejects
-mcpu=cortex-m52 outright. Support arrives in GCC 14.

---------

Co-authored-by: Frédéric Desbiens <frederic.desbiens@eclipse-foundation.org>
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-16 08:29:46 -04:00
Frédéric Desbiens 09c190a71d Added a CMake target for the Linux sample program (#622)
The CMake build produced libthreadx.a and nothing else, so trying ThreadX on
Linux meant using the Makefile beside the port instead. Build the demo the
Makefile builds, for the linux port and its SMP counterpart.

The target is behind an option that defaults off, so an ordinary build is
unchanged and still produces just the library. -DTHREADX_SAMPLE=ON adds it:

    cmake -S . -B build -DTHREADX_ARCH=linux -DTHREADX_TOOLCHAIN=gnu \
          -DTHREADX_SAMPLE=ON
    cmake --build build --target sample_threadx

The include path uses TX_COMMON_DIR rather than naming common or common_smp,
since the top level already resolves which of the two applies.

Verified by building and running both variants. Non-SMP prints

    **** ThreadX Linux Demonstration **** (c) 1996-2020 Microsoft Corporation

and SMP prints the SMP banner, both with the demo's thread counters advancing. A
default configure with no THREADX_SAMPLE has no sample_threadx target and still
produces libthreadx.a, so nothing existing moves.

Derived from the two example_build files in #404 by Yanfeng Liu, which had the
same goal. That change also rewrote the top level's SMP selection, added
common_smp/CMakeLists.txt and added ports_smp/linux/gnu/CMakeLists.txt; all three
have since arrived on dev by other routes, so only the sample targets were still
missing. The include path needed adjusting because the original depended on
THREADX_SMP being a string suffix, which it no longer is.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-15 17:28:33 -04:00
Frédéric Desbiens 1296cf1740 Corrected the S32Z280 data SRAM map against the Reference Manual and the part (#621)
The RTU has 1 MB of data SRAM in three contiguous banks, and this port
described it wrongly in both directions.

  DRAM0  0x31780000  256 KB  full core speed
  DRAM1  0x317C0000  256 KB  full core speed
  DRAM2  0x31800000  512 KB  half core speed

DRAM1 was not declared at all, so 256 KB of full-speed memory went unused
and the linker's DATA region stopped at 256 KB. DRAM2 was declared as
2 MB when it is 512 KB, which mattered more: the MPU mapped 1.5 MB past
the end of the bank, and that range aliases back onto its base. Anything
placed above 0x31880000 would have shared storage with the bottom of the
region silently -- no fault, two objects at one address. Nothing was
placed there yet, so this was a trap rather than a live defect.

Sources: S32Z2 Reference Manual Rev. 5, section 6.3.6 and Table 13, and
the board. Writing distinct values to all three banks and reading them
back shows 1 MB of independent storage, and the first word past DRAM2
returns the value written to its base, which is what fixes the size.

Note that NXP's own debugger memory map, s32z2e2_memory_regions.py in
S32 Design Studio, calls the last bank 2 MB. The Reference Manual and the
silicon agree it is 512 KB.

Also corrected the description of these banks throughout. They are all
RTU-local; the earlier comments treated locality as the thing that
distinguishes them, when the actual difference is clock speed. That is
why the cache benchmark uses DRAM2 -- caching a bank that already runs at
core speed shows nothing, which is a real effect the old wording
explained with the wrong cause.

Verified on the S32Z280-594EVB: builds clean, and the boot probes pass
six of six with the MPU, GIC, interrupts, caches and both protection
faults exercised.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-15 17:02:01 -04:00
Frédéric Desbiens 358e7a9ea5 Stopped the remaining example scripts naming archives that were deleted (#620)
#618 fixed the arm9 and arm11 shell scripts, but not their .bat counterparts,
and did not look at cortex_r4 and cortex_r5 at all. An audit of every build
script against the files actually present in its directory found the rest.

The .bat scripts for arm9, arm11, cortex_r4 and cortex_r5 still linked libc.a,
libgcc.a and, for arm11, libnosys.a from their own example_build directories,
and the cortex_r4 and cortex_r5 shell scripts did too. Those archives went in
6.1.10 under "Removal of unneeded files", so on Windows all four examples failed
exactly as the shell versions did before #618, and on Linux the two R-profile
ones still did.

Link through the compiler driver, as the other examples have since #594.

This does not make the examples link, and the change stops there deliberately.
All four now fail the same way,

    undefined reference to `_fini'

because their linker scripts define the .init and .fini sections but not the
_init and _fini symbols, which live in crti.o and crtn.o and are omitted by
-nostartfiles. Reviving four very old cores is separate work.

Correct the comment on EXAMPLES_EXPECTED_TO_FAIL again. It had cortex_r4 and
cortex_r5 failing for want of newlib multilib variants; they do not. All four
share the single cause above, and the multilib explanation was wrong for the
R-profile pair just as it was for arm9 and arm11.

Verified by running each shell script: cortex_r4 and cortex_r5 fail on _fini
rather than on missing files, matching arm9 and arm11. The .bat changes mirror
link lines proven that way in the same directories; they cannot be run here.

Three scripts are left alone and reported instead, because a blind edit could not
be verified: ports_module/cortex_m3 and cortex_m4 have Windows-only .bat scripts
that also name sources which are absent or differ in case, and
ports_smp/mips32_interaptiv_smp needs a MIPS toolchain.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-15 16:58:59 -04:00
Frédéric Desbiens 20d4b977f2 Removed generated build output and per-user IDE state, and fixed two stale link lines (#618)
* Removed generated build output and stopped two scripts naming deleted archives

Three kinds of file in the tree are produced by a build rather than written by
hand, and one pair of scripts still links archives that were deleted years ago.

Keil writes ThreadX_Library.plg on every build; the two committed copies are
HTML build logs from someone's machine. Code Composer generates the makefiles
under ports/c667x/ccs/example_build/tx/Release from the project files beside
them, so makefile, objects.mk, sources.mk, subdir_rules.mk, subdir_vars.mk and
ccsObjs.opt are all regenerated output.

The arm9 and arm11 sample builds link libc.a, libgcc.a and, for arm11,
libnosys.a from their own example_build directories. Those archives were removed
in 6.1.10 under "Removal of unneeded files", and libnosys.a in #594, but the link
lines were never updated, so both examples fail immediately with

    arm-none-eabi-ld: cannot find libc.a: No such file or directory

Link through the compiler driver instead, the shape every other example in the
tree uses since #594: the driver supplies libc and libgcc, and SYSCALL_LIB is
already defined in both scripts.

That does not make either example link, and the fix stops short of that on
purpose. With the archives no longer named, both now fail on

    undefined reference to `_fini'

because their linker scripts define the .init and .fini sections but not the
_init and _fini symbols, which live in crti.o and crtn.o and are omitted by
-nostartfiles. Making those two old cores build is a separate question from
removing a stale reference, so they stay in EXAMPLES_EXPECTED_TO_FAIL, with the
comment there corrected: it blamed newlib multilib packaging, which is true of
cortex_r4 and cortex_r5 but was never the reason for arm9 and arm11.

Reproduced throughout with arm-none-eabi-gcc 13.2.1.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Removed the Keil per-user state files and ignored them

Every Keil project in the tree carried a second file holding per-user state:
25 .uvoptx beside the 25 .uvprojx, 9 .uvopt beside the 9 .uvproj, and 3 .uvgui
multi-project workspace files. uVision rewrites all of them whenever a project is
opened, so they record whoever last had it open rather than anything about the
port: debugger selection, breakpoints, watch windows, window geometry.

Nothing in the tree references them, and every affected directory keeps its
.uvprojx or .uvproj, which is the file that actually describes the project.

Add ignore rules so they do not come back the next time someone opens a project
and commits.

1.4 MB across 37 files.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-15 16:40:19 -04:00
Frédéric Desbiens b1a48824ea Enabled ATCM on S32Z280, and left the other two banks off for a measured reason (#617)
entry.S programs ATCM to 0x30000000 and enables it at both exception levels. This
happens at EL2 deliberately: writing ENABLEEL2 from EL1 is silently ignored, which
was measured on BTCM and CTCM -- the base took and ENABLEEL10 took while
ENABLEEL2 stayed clear -- and the same write from EL2 sticks. tcm_enable() keeps
ENABLEEL10 as its success criterion for the same reason, since requiring both
would report failure for a bank that is usable at the level the caller runs at.

Programming ATCM also moves it off address 0, where CFGTCMBOOTx leaves it, so a
null-pointer write now faults instead of quietly landing in tightly-coupled
memory.

The boot image verifies rather than programs, preloads ATCM because ECC is enabled
and the check bits are not initialised by the core, and proves the bank holds data
both before the MPU is enabled and after. One MPU region covers it: the TRM
requires a region before an enabled TCM can be used, and an enabled TCM always
behaves as Non-cacheable Non-shareable Normal memory whatever the region says, so
only the permissions there matter.

BTCM and CTCM are left disabled, and that is a measurement rather than caution.
Enabling any second bank removes all measurable data-cache benefit:

    ATCM only                      cache gain 24%
    ATCM + BTCM                    cache gain  0%
    ATCM + BTCM + CTCM             cache gain  0%
    ATCM + BTCM at another base    cache gain  0%
    ATCM + CTCM, BTCM disabled     cache gain  0%

Five configurations, one variable. Not a particular bank, not its address, and not
the ECC preload: enabling a second bank at all. The benchmark buffer is in
non-RTU-local SRAM at 0x31800000, outside every TCM window, and CCSIDR reports the
same 16KB four-way cache throughout. I was wrong twice while narrowing this --
first blaming the preload, then blaming BTCM specifically -- and each was settled
by a run rather than by argument.

No erratum covers it. Checked the Cortex-R52 errata notice SDEN-857344 issue 19,
all twenty-five entries, and the S32Z2 0P91J mask set errata, whose RTU and R52
entries are ERR050509, ERR051107, ERR051153, ERR051441, ERR051613, ERR051614 and
ERR052126. A RAM pool shared between the RTU's last-level cache and the TCMs would
explain it, the LLC being documented as allocating ways to specific domains, but
that is a guess and it belongs with the other questions for NXP.

Little is lost meanwhile. The reference manual describes TCM_A as the bank
"optimized for small, regularly executed code such as interrupt service routines
or OS kernels", which is what a TCM is wanted for here, and enabling the other two
is one line each in entry.S once there is an answer.

Two checks were also wrong and are fixed. T5 reported every bank accessible while
two were disabled, because a disabled TCM's address range is serviced through AXIM
and memory answering there says nothing about the TCM; it now requires the bank to
be enabled as well. And T2 called tcm_enable() from EL1 for all three banks, which
would have switched on the very banks entry.S leaves off.

Verified on S32Z280 silicon: ATCM enabled and holding data before and after the
MPU, the cache benchmark back to 24%, protection checks X2 and X4 unchanged, and
the ThreadX demo still reporting 100 ticks with 20 preemptions.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-14 16:14:24 -04:00
Frédéric Desbiens 27891f3b84 Read the TCM configuration off the part, and corrected what we thought we knew (#615)
The BSP has recorded since bring-up that the TCMs are inaccessible at reset
"because nothing has programmed the TCM region registers yet". Reading those
registers shows the reason was wrong.

    ATCM  0x0000011F   64KB, 1 wait state, ENABLED at EL2 and EL1/EL0
    BTCM  0x00000014   16KB, 0 wait states, disabled
    CTCM  0x00000114   16KB, 1 wait state, disabled

Every BASEADDRESS field is zero, so ATCM is live as 64KB at address 0x00000000,
not at the 0x30000000 the reference manual documents. The reads that faulted were
of an address the TCM is not at. ATCM's enables reset set because CFGTCMBOOTx is
tied high on this part, which the Cortex-R52 TRM gives as the one exception to
"at reset all bits are 0 apart from SIZE and WAITSTATES". BTCM and CTCM really
are disabled.

The sizes and wait states match the S32Z2 reference manual exactly -- TCMA 64KB
with one wait state, TCMB 16KB with none, TCMC 16KB with one -- so the TRM's field
layout, NXP's documented configuration and the silicon all agree. That agreement
is the point of reading before writing.

ECC is implemented and enabled: IMP_MEMPROTCTLR reads 0x00000011, both RAMPROTIMP
and RAMPROTEN set. TRM 6.2.2 therefore applies rather than being hypothetical: a
TCM location must be written before it is read, or the read reports an error --
which looks exactly like "the TCM is not accessible" and sends the reader back to
region registers that were already correct. The preload widths differ too, ATCM
needing 64-bit aligned STRD or STM where BTCM and CTCM accept 32-bit stores, so a
C loop over unsigned int would leave ATCM's check bits invalid.

tcm.c reads and decodes only; nothing is programmed here. The layouts in tcm.h are
quoted from TRM r1p3 section 3.3.94 table 3-136 and section 3.3.76 table 3-114,
not inferred from a neighbouring register: BASEADDRESS is [31:13] where
IMP_PERIPHPREGIONR uses [31:12], and assuming the analogy would have been wrong by
one bit in the same way the PRBAR shift was.

The boot image reports all of it and flags any size that disagrees with the
reference manual, so a part configured differently says so rather than being
silently assumed to match this one.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-14 14:18:50 -04:00
Frédéric Desbiens ea91beae42 Verified nested FIQ handling on S32Z280 silicon (#614)
#613 exercised the FIQ nesting routines on the FVP. What the model could not show
is whether a real GIC-600 routes Group 0 to FIQ the same way, which is the reason
to run it here. It does, and the counts match the model exactly.

The Group 0 support ports across unchanged: IGRPEN0 and BPR0 on the CPU interface,
Group 0 in the distributor, gicv3_enable_sgi_group0, gicv3_send_sgi_group0 through
ICC_SGI0R, and the separate Group 0 acknowledge and end-of-interrupt pair. entry.S
routes the EL1 FIQ vector into _tx_thread_fiq_context_save with the acknowledge
before nesting starts, and leaves FIQ unmasked on the drop to EL1 when FIQ support
is compiled in, for the same reason as on the FVP: tx_thread_stack_build only
clears a thread's F bit in that configuration.

One structural difference from the FVP cost a link. This example reports faults
through FAULT_TAIL rather than FAULT_REPORT, so the vector table needed a new
el1_fiq_entry label that falls back to fault_el1_fiq. Placing that label inside the
TX_R52_USE_THREADX_IRQ guard broke s32z280_boot.elf, which does not define it: the
vector reference is unconditional, so the label has to be too. It now sits outside
the guard and carries its own, the same shape the demo_m2 link break in #613
forced on the FVP side.

Verified on S32Z280 silicon:

    F1 FIQ delivered and dispatched            PASS
    low-priority FIQ count  = 0x00000015       21
    high-priority FIQ count = 0x00000014       20
    nested FIQ count        = 0x00000014       20 of 20 nested
    max FIQ depth           = 0x00000002
    FIQ depth now           = 0x00000000
    F2 FIQ nested inside an FIQ handler        PASS
    F3 FIQ nesting unwound to depth zero       PASS
    F4 IRQ tick undisturbed by FIQ work        PASS
    F5 lower-priority thread still scheduled   PASS
    F6 no unexpected Group 0 INTID             PASS

No regression, both configurations checked on the board. In the FIQ build the
ThreadX demo still reports 100 ticks with 20 preemptions and the boot image still
passes its cache and protection checks. In the default build the demo is unchanged
and the image links no FIQ or Group 0 symbol at all, so the work is absent rather
than dormant where it is not wanted -- worth confirming on hardware rather than
reasoning about, because the default build now reaches the FIQ vector through a new
label even though that label only branches to the fault reporter.

entry.S assembles in all four combinations of TX_R52_USE_THREADX_IRQ and FIQ
support, on both toolchains, and every file builds with GNU without warnings and
with Arm Toolchain for Embedded 22.1.0.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-14 13:34:34 -04:00
Frédéric Desbiens 19d90a49c4 Exercised the nested FIQ path, the last pair nothing had ever called (#613)
_tx_thread_fiq_nesting_start and _tx_thread_fiq_nesting_end complete the set:
after #611 and #612 covered IRQ nesting on the model and on silicon, these two
were the remaining routines compiled into every build and entered by nothing.
demo_fiq.elf enters them.

FIQ needs more of the GIC than IRQ does. With a single security state the
controller delivers Group 0 as FIQ and Group 1 as IRQ, so an interrupt only
arrives as an FIQ if it has been moved into Group 0, the distributor and the CPU
interface both have Group 0 enabled, and it is acknowledged through the Group 0
registers. Group 1's acknowledge returns the spurious INTID for a Group 0
interrupt and leaves it pending, which would present as a storm rather than as an
error. gicv3.c gains IGRPEN0, BPR0, gicv3_enable_sgi_group0,
gicv3_send_sgi_group0 and the Group 0 acknowledge and EOI pair. ICC_SGI0R differs
from ICC_SGI1R only in opc1, 2 against 0, and each raises into its own group.

entry.S routes the EL1 FIQ vector into _tx_thread_fiq_context_save with the same
ordering the IRQ path needed: acknowledge in FIQ mode before nesting starts, then
nesting_start, service, nesting_end, and end-of-interrupt last. It also leaves
FIQ unmasked on the drop to EL1 when FIQ support is compiled in, because
tx_thread_stack_build only clears a thread's F bit in that configuration and
nothing else ever clears it, so an FIQ raised before the first thread ran would
otherwise be silently ignored.

Nesting an FIQ means taking an FIQ while an FIQ handler runs, which one source
cannot show, so two Group 0 SGIs are used with the second at a numerically lower
priority. The low one's handler raises the high one.

One mistake worth recording, because the guard it needed is not obvious.
TX_ENABLE_FIQ_SUPPORT is PUBLIC on the threadx target, so it reaches every image
as soon as the library is built with FIQ -- including images that link no
interrupt controller at all. demo_m2 is one of those: no gicv3.c, no
irq_dispatch.c. Referencing gicv3_acknowledge_group0 and board_fiq_service from
the FIQ vector broke its link outright. The vector is now gated on
TX_R52_USE_THREADX_IRQ as well, which is how the IRQ vector has always been
gated, and images without the infrastructure keep the fault reporter.

Verified on FVP_BaseR_AEMv8R. F1 covers the whole Group 0 chain on its own --
IGRPEN0, the ICC_SGI0R encoding, IAR0 and EOIR0, the EL1 vector and F being
unmasked -- so a failure there points somewhere other than the nesting routines.
The demo reports 21 low-priority FIQs, 20 high-priority, 20 of them nested, max
depth 2, depth unwound to zero, the IRQ tick undisturbed, threads still
scheduled and no unexpected Group 0 INTID.

All six images pass in the FIQ configuration, and the default build links no FIQ
or Group 0 symbol at all, so the work is absent rather than dormant where it is
not wanted. Both toolchains build every file, GNU with no warnings and Arm
Toolchain for Embedded 22.1.0 as well.

Not done here: the S32Z280 example is untouched, so FIQ on silicon is a separate
change, as IRQ nesting was.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-14 11:42:59 -04:00
Frédéric Desbiens 5a7f4f8f7c Verified nested IRQ handling on S32Z280 silicon (#612)
#611 exercised the nesting routines on the FVP and established what breaks them:
the interrupt must be acknowledged before nesting starts, or the still-pending
level-asserted timer is retaken the moment IRQ is enabled and recurses until the
stacks are gone. entry.S here carries the same ordering for the same reason, and
irq_dispatch.c splits the same way, board_irq_service taking an
already-acknowledged INTID while board_irq_handler keeps its old shape as the
non-nesting entry point.

What the model could not answer is whether a real GIC-600 agrees, and two things
could have differed.

The first is the number of implemented priority bits. Equal priorities do not
preempt and it is the low bits that vanish, so if the timer and the SGI collapse
to one value after truncation then nesting cannot happen at all -- and the test
would fail without saying why. gicv3_priority_bits discovers the count by writing
0xFF to a priority byte and reading back which bits stick, board_init records the
two effective values, and check P1 requires the SGI to still outrank the timer.
This silicon keeps five bits, the same as the FVP, so 0xA0 and 0x50 stay distinct;
that is now measured and reported rather than assumed.

The second is whether an SGI raised on real hardware is delivered at all. ICC_SGI1R
is a 64-bit AArch32 CP15 register whose encoding does not transcribe from the
AArch64 alias, so check N1 raises one from thread context and requires delivery
before nesting is involved. It arrives.

Verified on S32Z280 silicon:

    priority bits   = 0x00000005
    timer effective = 0x000000A0    sgi effective = 0x00000050
    P1 SGI outranks timer after truncation     PASS
    N1 SGI delivered and dispatched            PASS
    max depth = 0x00000002   nested SGIs = 0x00000032   depth now = 0x00000000
    N2 SGI nested inside another handler       PASS
    N3 nesting unwound to depth zero           PASS
    N4 tick still advancing after nesting      PASS
    N5 lower-priority thread still scheduled   PASS
    N6 no spurious or unexpected interrupts    PASS

Fifty nested SGIs across fifty ticks, one per tick, and depth never exceeded two.

No regression, checked in both configurations on the board. In the nesting build
the ThreadX demo still reports 100 ticks with 20 preemptions and the boot image
still passes its cache and protection checks. In the default build the image links
no nesting symbols at all and the demo is unchanged. Both toolchains compile every
file, GNU with no warnings and Arm Toolchain for Embedded 22.1.0 as well.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-14 11:07:06 -04:00
Frédéric Desbiens 559460bed9 Exercised the nested IRQ path, which had never once been entered (#611)
_tx_thread_irq_nesting_start and _tx_thread_irq_nesting_end have shipped in this
port since it was written, compiled into every build, and nothing had ever called
either one. Not on the model, not on silicon, not in any demo. demo_nesting.elf
enters them.

Provoking nesting needs two sources with different priorities. The generic timer
PPI was already there; the second is an SGI, which a core can raise on itself.
gicv3.c gains gicv3_enable_sgi and gicv3_send_sgi for that. ICC_SGI1R is 64-bit,
so in AArch32 it is an MCRR rather than an MCR, and the AArch64 name the
Cortex-A72 example uses, S3_0_C12_C11_5, does not transcribe to the AArch32 CP15
space. The encoding here is confirmed by check N1 in the demo, which raises an SGI
from thread context and requires it to be delivered and dispatched.

The order of the pairing is the whole difficulty, and getting it wrong does not
fail gracefully. The interrupt must be acknowledged BEFORE nesting starts. Reading
ICC_IAR1 is what raises the GIC running priority to this interrupt's own, which
masks it and everything of equal or lower priority; only then is re-enabling IRQ
safe. My first attempt called nesting_start first and acknowledged inside the
handler, so the still-pending, still-level-asserted timer was taken again the
instant IRQ was enabled, and again, until the IRQ and System stacks were
destroyed. It presented as garbage on the console and a hang with no fault to
point at, and it broke demo_m3 and demo_threadx while leaving boot_check, demo_m2
and demo_mpu passing, because only the first two depend on the tick advancing. The
Cortex-R5 example BSP states the requirement in one line: "ensure all IRQ
interrupts are cleared prior to enabling nested IRQ interrupts."

So entry.S now acknowledges in IRQ mode, carries the INTID in r4 -- which survives
the mode switch, since only SP and LR are banked -- and also pushes it on the IRQ
stack so a nested level reusing r4 cannot lose the outer level's value.
End-of-interrupt waits until after nesting_end, in IRQ mode with interrupts
masked, so dropping the running priority cannot re-admit the same interrupt.

board_irq_handler splits in two. board_irq_service does the middle part on an
already-acknowledged INTID and neither acknowledges nor EOIs; board_irq_handler
keeps its old shape as the non-nesting entry point, so images built without
TX_ENABLE_IRQ_NESTING behave exactly as before.

The nesting instrumentation in irq_dispatch.c is inert unless an image asks for
it. board_nest_provoke gates the SGI that the timer handler raises, so every other
image sees the handler it always had.

Verified on FVP_BaseR_AEMv8R. The nesting demo reports max depth 2, fifty nested
SGIs across fifty ticks, depth unwound to zero, the tick still advancing, the
low-priority thread still scheduled, and no spurious or unexpected INTIDs. In the
same nesting configuration the five existing images all pass, so the split did not
disturb the ordinary path. Both toolchains build it, GNU with no warnings and Arm
Toolchain for Embedded 22.1.0 as well.

Not done here: the S32Z280 example keeps its own irq_dispatch.c and entry.S and is
untouched, so nesting on silicon is a separate change.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-13 17:13:34 -04:00
Frédéric Desbiens b877305d13 Verified lazy VFP context switching on S32Z280 silicon (#609)
The FVP proved the lazy save and restore path at AR1/M5, but the board had never
run it. Doing so needed one thing the model did not: the FPU has to be turned on.

entry.S opens CPACR for CP10/CP11 and sets FPEXC.EN at EL1, guarded by __ARM_FP.
EL2 already cleared HCPTR.TCP10/TCP11, but both of the EL1 gates read 0 out of
reset, so a floating-point instruction raised an Undefined Instruction exception
before this. tx_thread_vfp_enable() does not help: it sets the per-thread software
flag that makes the context switch save and restore the registers, and never
touches the hardware. Enabling it is the BSP's job. The __ARM_FP guard is what
keeps a soft-float build assemblable, since "vmsr fpexc" is not a valid
instruction for that target at all.

demo_vfp_s32z280.c is the FVP's demo_m5.c test design over the LINFlexD console.
The design is kept deliberately, because the two halves of the VFP context path
need separate provocation: "fp check" holds eight live doubles across
tx_thread_sleep, eight being enough to force the callee-saved D8-D15 bank that a
solicited switch must preserve, while "fp busy" sits at the lowest priority and
never sleeps, so the tick interrupts it mid-computation and exercises the
interrupt half, D0-D15 plus FPSCR. Only these two threads opt in, so the test
also shows that opting in is what does the work.

s32z280_vfp.elf is gated on TX_R52_ENABLE_VFP, as demo_m5.elf is in the FVP
example, since the image is meaningless unless the library was built with a
floating-point ABI.

Also made this example's -Wl,--no-warn-rwx-segments conditional on the compiler
being GNU. #604 did that for the FVP example and this file still carried the
literal, so ld.lld failed the link with "unknown argument". With that fixed the
S32Z280 images build with Arm Toolchain for Embedded too.

Verified on S32Z280 silicon:

    V2 D8-D15 bank preserved across switches   PASS   50 solicited switches
    iterations = 0x000140A8                           82,088 interrupted rounds
    corruptions = 0x00000000
    V3 interrupted FP thread made progress     PASS
    V4 no FP corruption across interrupts      PASS
    filex_ptr = 0xF11EF11E
    V5 VFP flag did not alias filex_ptr        PASS
    PASS lazy VFP context switch verified on silicon

No regression: the soft-float build is warning-free and its image contains no vmsr
at all, confirming the guard elides the block, and the existing demo still reports
100 ticks with 20 preemptions on the board. Both S32Z280 images also build with
Arm Toolchain for Embedded 22.1.0, the hard-float VFP image included. No FVP file
is touched.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-13 17:07:36 -04:00
Frédéric Desbiens acdc02b2fc Assembled the code behind feature macros, and fixed the POP it found (#608)
scripts/check_clang.sh assembled every source with default flags, so the
preprocessor discarded each #ifdef block before the assembler saw it. Nothing in
the tree had ever assembled a guarded path. That covers the VFP context save and
restore in ten ports, and 218 files carrying TX_LOW_POWER or
TX_ENABLE_EXECUTION_CHANGE_NOTIFY.

Turning those on found a defect. The Cortex-M0 and Cortex-M23 execution-profile
paths bracket their call with

    PUSH    {r0, lr}
    BL      _tx_execution_isr_enter
    POP     {r0, lr}

and the last of those is invalid on Armv6-M and Armv8-M Baseline, where the
16-bit Thumb POP takes r0-r7 and pc and nothing else. GNU rejects it as well --
"cannot honor width suffix" -- so TX_ENABLE_EXECUTION_CHANGE_NOTIFY and
TX_EXECUTION_PROFILE_ENABLE have never been buildable on either port with either
toolchain. Four files, all the same shape.

The fix pops into a scratch register and moves it, MOV to a high register being
permitted where POP is not. r1 is free: the BL may clobber r0-r3, which is the
reason r0 is saved in the first place. Disassembling the result gives
push {r0, lr} / bl / pop {r0, r1} / mov lr, r1 / bx lr, one 16-bit instruction
more than before and otherwise the same.

Two findings that were not defects, recorded in the script so they are not
rediscovered:

Cortex-R4 needs an -mfpu to assemble its VFP path, because its FPU is an option
rather than part of the core. GNU fails identically without one, so this is a
flags requirement and not a toolchain divergence.

The A profile ports must not be given one. Adding -mfpu=vfpv3-d16 uniformly broke
28 files with "register expected", because those ports save D16-D31 and a -d16
FPU does not have those registers. Their defaults were already right.

The new stage runs under --asm-only as well, needing no target C library, and
reports 37 of 37 VFP files, 8 of 8 TX_LOW_POWER and 218 of 218
TX_ENABLE_EXECUTION_CHANGE_NOTIFY. Restoring the POP for one run makes it fail
with 217 of 218 and name the file and the error, so the stage is not vacuous.
The other four stages are unchanged: 711 of 711 assembled, 185 of 185 common C
sources for each of nine cores, 42 of 42 script-driven examples and 5 of 5 CMake
images.

The fixed code is verified to assemble with both toolchains and to encode as
intended. It is not verified running: there is no Cortex-M0 or Cortex-M23 model
here, and these are context save and restore paths, so that gap is worth stating.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-13 10:14:53 -04:00
Frédéric Desbiens 010a6c9fbb Gave the gnu ports a CMake build, which most of them lacked (#607)
The project guidelines ask for CMake and Ninja, but only 15 of the 59 gnu port
directories had a CMakeLists.txt. None of the 27 AArch64 ports had one, so the
architecture whose examples were repaired over the last few changes still could
not be built the way the project says to build it, and nothing in CI could
compile it.

Add a CMakeLists.txt to the 44 that lacked one. Three of them are templates in
ports_arch, because 34 of the 44 are generated: the ARMv7-A and AArch64 source
lists are uniform within each family, so one template per family serves every
core in it and update.sh distributes it. The other 10 ports have no generator
and get their own file.

Add the toolchain files those ports select, following the shape of
cmake/cortex_a9.cmake. AArch64 needs a base file of its own rather than a
variant of arm-none-eabi.cmake: it has no -marm or -mthumb to choose between and
no -mfloat-abi, and aarch64-none-elf-gcc rejects -mlong-calls outright, so that
flag cannot be carried across. The tools are named without a path, unlike
cmake/cortex_r52.cmake which pins one, because pinning 30 files to a single
machine's directory layout is the problem the previous change removed from the
launch configurations.

Three toolchain files cover ports that already had a CMakeLists.txt but no way
to select it: the Armv8-M mainline gnu ports, cortex_m33, cortex_m55 and
cortex_m85. Without cmake/<arch>.cmake the documented invocation cannot reach
them.

The top level needed one fix. It derives the SMP port directory as
<arch>_smp, but ports_smp/linux and ports_smp/win64 predate that convention and
carry no suffix, so those two could never be configured. Fall back to the bare
name when the suffixed directory is absent. The check only fires when the
suffixed directory does not exist, so no port that already resolved changes
behaviour, and ports_smp/win64's existing CMakeLists.txt becomes reachable too.

Verified by configuring and building every one: 53 of 53 static libraries build
with cmake -G Ninja, using Arm GNU Toolchain 14.3.Rel1 for both arm-none-eabi
and aarch64-none-elf. That covers the 44 new ports plus the 9 that already
worked, and includes ports_smp/linux, which failed before the fallback.
scripts/check_ports.sh passes, so the three templates and their 34 generated
copies agree.

Six gnu ports are still outside the CMake build, all for want of a compiler
rather than a CMakeLists.txt: rxv1, rxv2 and rxv3 need the Renesas RX GNU
toolchain and mips32_interaptiv_smp needs a MIPS one, neither of which is
available here, so writing toolchain files for them would mean shipping
untested guesses. risc-v32 and risc-v64 already build through their own
differently named toolchain files.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-12 12:43:57 -04:00