mirror of
https://github.com/eclipse-threadx/threadx.git
synced 2026-10-06 06:59:08 +08:00
515ab8aba1a42a9d85e18288e1a76e579b177831
445
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
515ab8aba1 |
Added the missing memory barrier so ARMv7-A SMP schedulers no longer miss preemptions (#704)
_tx_thread_schedule stores the newly selected thread into _tx_thread_current_ptr[core] and then reloads _tx_thread_execute_ptr[core] to detect a concurrent scheduling decision made by another core. On the other side, _tx_thread_smp_core_interrupt stores the execute pointer and then reads the current pointer to decide whether an inter-core interrupt is required. This is a store-buffer pattern: without a barrier both cores can observe stale values, the interrupt is skipped, and a ready thread with the highest priority is never scheduled. The barrier was added to the ARMv8-A SMP scheduler in 6.2.1, but the ARMv7-A SMP ports were left untouched even though they implement the same protocol. This adds the corresponding DMB to the Cortex-A5, Cortex-A7 and Cortex-A9 SMP schedulers for both the AC5 and GNU toolchains. The GNU variants were verified by assembling them with arm-none-eabi-gcc 13.2.1 for their respective cores. Refs #209 Assisted-by: Copilot (Opus 5) <noreply@github.com> |
||
|
|
e99f0d4207 |
Corrected the module kernel stack size so it no longer overstates the usable stack (#701)
The module manager recorded tx_thread_module_kernel_stack_size as the raw TXM_MODULE_KERNEL_STACK_SIZE constant, but the end of the kernel stack is aligned downwards to an eight-byte boundary while _txm_module_manager_object_allocate only guarantees ULONG alignment. The recorded size could therefore overstate the usable stack by up to seven bytes. The scheduler copies this value into tx_thread_stack_size whenever a user mode module thread enters the kernel, so the overstated value is visible to RTOS-aware debuggers and to anything built on it. The size is now derived from the aligned end minus the start. Also documented that TX_ENABLE_STACK_CHECKING is not supported for module threads. Refs #181 Assisted-by: Copilot (Opus 5) <noreply@github.com> |
||
|
|
945f5f5caa |
Moved the ARC ISR enter callout onto the system stack (#700)
_tx_thread_context_save() calls _tx_execution_isr_enter when TX_ENABLE_EXECUTION_CHANGE_NOTIFY is defined. On the path where an interrupt preempted a running thread, that call was made before the switch to the system stack, so the 32-byte call frame and the whole stack footprint of the callout were taken from the interrupted thread's stack, on top of the 160-byte interrupt frame the port had already allocated there. The callout is supplied by the application, so its stack usage is not bounded by ThreadX, and it is charged to every thread that happens to be running when an interrupt arrives. The two other callout sites in the same routine, the nested save and the idle system save, already run on the system stack, as does the _tx_execution_isr_exit call in _tx_thread_context_restore. _tx_thread_schedule was reordered in 6.1.9 so that _tx_execution_thread_enter runs on the system stack rather than the thread stack; the same reorder was never applied to the context save. The switch to the system stack now happens before the callout in the ARCv2_EM, ARC_HS and SMP ARC_HS ports. _tx_thread_context_fast_save is unchanged because the fast interrupt path never switches stacks by design. Fixes #149 Assisted-by: Copilot (Opus 5) <noreply@github.com> |
||
|
|
29afcc3946 |
Fixed a kernel stack leak when deleting user-mode module threads (#692)
* modules: free kernel stack on thread deletion Signed-off-by: Prashit Vora <prashitvora2006@gmail.com> * Preserved the thread object release when the kernel stack cannot be freed Releasing the kernel stack ahead of the thread object made a failure of the kernel stack deallocation abort the thread object release. The thread had already been deleted at that point, so the thread object would have stayed allocated for the lifetime of the module. The thread object is now always released once the delete succeeds, and the kernel stack failure is reported only when it does not mask a thread object failure. --------- Signed-off-by: Prashit Vora <prashitvora2006@gmail.com> Co-authored-by: Frédéric Desbiens <frederic.desbiens@eclipse-foundation.org> Assisted-by: Copilot (Opus 5) <noreply@github.com> |
||
|
|
b0ec8bfbb9 |
Fixed the invalid module data pointers in the absolute module load (#699)
An absolutely located module has its code and its data placed at two independent fixed addresses by the module's linker script. The module preamble carries the code and data sizes but not the data address, so _txm_module_manager_absolute_load() could not determine where the module's data area was. It computed txm_module_instance_data_start from the code size and the preamble size, which yields a size rather than an address, and it set txm_module_instance_module_data_base_address one past the end of the byte pool allocation. Added _txm_module_manager_absolute_load_extended(), which accepts the module's data area address from the caller. Deprecated _txm_module_manager_absolute_load(), which now forwards to the extended service with an unknown data area location and rejects modules that request memory protection, since the memory protection hardware cannot be programmed to cover an unknown data area. Fixes #450 Assisted-by: Copilot (Opus 5) <noreply@github.com> |
||
|
|
d7789f0b12 |
Fixed the clobbered return address in the RISC-V context save (#696)
_tx_thread_context_save() returns to its caller with ret, which uses the return address held in ra. When TX_ENABLE_EXECUTION_CHANGE_NOTIFY was defined, the call to _tx_execution_isr_enter overwrote ra with the address of the instruction following the call, so the subsequent ret returned into _tx_thread_context_save itself instead of the interrupt service routine. The return address is now saved on the stack around the call and recovered afterwards, which is the same idiom already used by the Arm ports. The fix covers all three affected paths (nested save, thread save and idle system save) in the risc-v32 GNU, risc-v32 IAR and risc-v64 GNU ports. Fixes #348 Assisted-by: Copilot (Opus 5) <noreply@github.com> |
||
|
|
508af549da |
Fixed the incorrect loop bound constant in the IAR file lock support (#695)
The IAR multithreaded library support code allocates its file lock mutexes from an array of _MAX_FLOCK entries, but the wrap-around check and the exhaustion check in __iar_file_Mtxinit() both compared against _MAX_LOCK, the bound of the unrelated system lock mutex array. When _MAX_FLOCK is greater than _MAX_LOCK, the free mutex index wrapped early and the exhaustion check reported failure while free entries remained, so *m was set to TX_NULL and the application faulted the first time a file lock was taken. When _MAX_FLOCK is smaller than _MAX_LOCK, the free mutex index was allowed to run past the end of __tx_iar_file_lock_mutexes and the exhaustion check could never fire. Corrected all four comparisons in each of the 27 copies of tx_iar.c. Fixes #444 Assisted-by: Copilot (Opus 5) <noreply@github.com> |
||
|
|
17ff21a07f |
Fixed the garbled comment in the POSIX condition variable sources (#694)
The comment above the internal semaphore lookup in the POSIX condition
variable implementation was garbled: it contained a capitalisation typo
("COndition") and an incomplete sentence ("into a semaphore a cast").
The four affected sources now share a single, correct wording.
This is a comment-only change; there is no functional impact.
Fixes #419
Assisted-by: Copilot (Opus 5) <noreply@github.com>
|
||
|
|
c469a7756b |
Fixed the missing immediate prefix on MOV in Cortex-M schedulers (#693)
The BASEPRI-masking path in tx_thread_schedule wrote "MOV r0, 0" rather than "MOV r0, #0". UAL requires the "#" prefix on an immediate operand. GNU as and the LLVM-based assemblers accept the unprefixed form and emit the intended encoding, but stricter assemblers reject it outright, so the affected ports could not be built with those toolchains. The GNU and AC6 sources had already been corrected; this brings the IAR and AC5 sources into line. Verified that both spellings assemble to the same Thumb-2 encoding (f04f 0000), so this is a source-correctness fix with no change in generated code or runtime behaviour. Covers 31 occurrences across the Cortex-M3, M4, M7, M33, M52, M55 and M85 ports, their module manager counterparts, and the shared ARMv7-M and ARMv8-M architecture sources. Fixes #461 Assisted-by: Copilot (Opus 5) <noreply@github.com> |
||
|
|
13c8c768c7 |
Added the boot-at-EL1 option to the S32Z280 entry path (#690)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The Armv8-R AEM FVP entry path has carried TX_R52_BOOT_AT_EL1 since it was written, for the case its own comment describes: "an earlier boot stage or a vendor EL2 monitor has already dropped privilege to EL1". This board's entry path did not, so a kernel could not be built as a guest on it at all -- and it is the board where that matters most, because it is the one with silicon behind it. The bracket is the whole change. Everything from the Thumb reset trampoline to the ERET goes inside #ifndef TX_R52_BOOT_AT_EL1, and the #else supplies a one-instruction A32 _start that branches to el1_entry. A32 AND NOT T32, which is the one real difference from the standalone entry. The core resets in Thumb state here because the RTU boot instruction NXP plants is a T32 branch, but a guest is not reached by reset: it is reached by the monitor's ERET, and the monitor chooses the state through SPSR.T. Get the two out of agreement and the guest dies on its first instruction with an undefined-instruction exception, which looks exactly like a bad entry address and sends the reader to the loader instead of to the ERET. WHAT THE MONITOR INHERITS is enumerated at the #ifndef, next to the code it replaces rather than in a document, because that is where somebody adding a third board will be looking. This board's EL2 block is considerably larger than the model's, and each item on the list is something a guest at EL1 provably cannot do rather than something it merely does not: CNTFRQ is writable only at the highest implemented exception level and reads zero out of reset; HCPTR.TCP10/TCP11 reset set, trapping every EL1 floating-point access; HSCTLR.TE is an EL2 register (SCTLR.TE is EL1's, and el1_entry still clears it below); ICC_HSRE.SRE makes every other ICC_* and ICH_* register exist at all; the low-latency peripheral port enables reset to zero and an EL1 write to that register traps to EL2; and the TCM enables are per-core with ENABLEEL2 SILENTLY IGNORED from EL1 -- measured on both BTCM and CTCM, the base took and bit 0 took while bit 1 stayed clear. CNTHCTL.PL1PCTEN and PL1PCEN are the deliberate omission from that list, and the note says why. This path opens both, because a standalone kernel owns the physical timer. A monitor that TIME-partitions its guests must not: a partition's physical time keeps running while it is descheduled, so a guest reading it can observe that it was not running. That is the monitor's decision rather than this file's, which is why the list says what a guest cannot do rather than what a monitor should. Verified both ways. The three standalone images build and link unchanged, and a kernel built with the option boots at EL1 on a S32Z280-594EVB under an EL2 monitor, runs two threads through a queue and a semaphore, and reports back -- with no other change to the kernel or to its port. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
8c681c188e |
Asserted that the Cortex-R52 port refuses the options it documents as refused (#687)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
ports/cortex_r52/gnu/CMakeLists.txt rejects two option combinations at
configure time -- TX_R52_ENABLE_VFP without TX_R52_FLOAT_ABI=hard, and
TX_R52_ENABLE_FIQ_NESTING without TX_R52_ENABLE_FIQ. Nothing exercised
either. A guard that has stopped firing looks exactly like a guard nobody
has tripped, so both could have been silently disabled by a typo in a
variable name at any point and no check would have noticed.
check_gcc.sh gains a sixth stage that configures each rejected combination
and asserts the configure fails. It needs no new harness: the script already
runs cmake as a subprocess for the CMake example builds, and the port's
CMakeLists is included by the toolchain file alone, so no example flags are
needed and the two negative configures stop almost immediately.
Two things the stage does that a thinner version would not.
It asserts the message text, not just the exit status. A configure that
fails for an unrelated reason would otherwise be recorded as a guard doing
its job.
It also configures the SUPPORTED combination and requires that to succeed.
Two negative assertions on their own are satisfied by a guard that rejects
everything -- the port would be unbuildable and the check would still pass.
The positive case is what separates a guard that works from one that is
merely always on.
Verified negatively, four deliberate breaks, each caught:
VFP guard condition forced false -> "configure succeeded, but this
combination cannot build"
FIQ nesting guard forced false -> the same, on that case
guard fires, message text changed -> "configure failed, but not on the
expected guard"
VFP guard condition forced true -> "the supported combination was
refused"
The fourth is the one the positive case exists for and the only one a
refusals-only stage would have missed. ports/cortex_r52/gnu/CMakeLists.txt
was confirmed byte-identical to dev afterwards.
Full run passes: exit 0, all six stages, and --asm-only correctly skips the
new one.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
57390fc0fe |
Refused the Cortex-R52 VFP option without a hard float ABI (#686)
TX_R52_ENABLE_VFP with the default soft float ABI cannot build. The option defines TX_ENABLE_VFP_SUPPORT, which enables the VMRS, VSTMDB and VLDMIA blocks in the port assembly, and -mfloat-abi=soft leaves the assembler with no FPU to accept them. The configuration fails with eight errors of the form tx_thread_system_return.S:118: Error: selected processor does not support `vmrs r4,FPSCR' in ARM mode none of which mentions the float ABI, so a user has to reason from VMRS back to the option that enabled it. The guard that exists said otherwise. It warned that "the compiler will not emit floating-point instructions, so the VFP context path will never be exercised", which describes a build that succeeds and is merely pointless -- and then let configure finish, so the warning scrolled past well before the assembler errors appeared. It is now a FATAL_ERROR that names the fix, which is what the same file already does six lines above for TX_R52_ENABLE_FIQ_NESTING without TX_R52_ENABLE_FIQ. That combination is rejected for being "meaningless", while this one, which cannot assemble at all, was only warned about. The severities were the wrong way round. The option's definition also moves below the check, so the block reads like the FIQ nesting one. The ABI is not promoted to hard automatically. TX_R52_FLOAT_ABI is a cache variable the user may have set deliberately, and silently overriding an explicit choice is worse than refusing a combination that cannot work. readme_threadx.txt carried the same claim, and its option list marked the FIQ nesting dependency inline but not this one. Both corrected. No regression test. Nothing in the tree asserts a configure-time failure -- there is no harness for it, and the sibling FIQ nesting guard has none either -- so a test for this would have to introduce that mechanism for one case. Verified by hand in both directions instead: the soft-ABI combination now stops at configure with the message above, and the hard-ABI feature build (VFP, FIQ, IRQ nesting, FIQ nesting) builds its eight images clean and passes ctest 8/8 on the Armv8-R AEM FVP. scripts/check_gcc.sh passes unchanged; it configures the Cortex-R52 CMake stage without TX_R52_ENABLE_VFP, so the new branch is not on its path. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
a2800fef16 |
Covered the SMP suspension teardown and the long byte pool search (#677)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
Against the merged SMP coverage report -- every build configuration instrumented and unioned -- sixty-four lines of common_smp/src were uncovered, 5114 of 5178. Fifty-three of them are closed here and the report reads 5167 of 5178. The SMP coverage floor goes from 98 to 99 with it. Thirty-six of the sixty-four were one loop repeated four times: the walk in tx_block_pool_delete, tx_byte_pool_delete, tx_event_flags_delete and tx_queue_delete that releases every thread suspended on the object with TX_DELETED. The suite deletes all four object types after every single test, and that is exactly why the loop never ran. test_control_cleanup in the ThreadX suite deletes the application's objects first and its threads last, so a test that ends with a thread parked on a queue has that thread walked out of it by tx_queue_delete. The SMP suite's cleanup deletes the threads first, deliberately -- it was changed so that no application-owned object is still referenced when the object loops run, which is what stopped a class of teardown hang. The side effect is that all four deletes now run against an empty suspension list. tx_semaphore_delete is the one member of the family that was already covered, because threadx_semaphore_delete_test deletes a busy semaphore on purpose. threadx_object_delete_suspension_test is the same idea for the other four. Two threads suspend on each of a block pool, a byte pool, an event flags group and a queue; the control thread waits on the object's own suspended count through tx_*_info_get rather than on an ordering it cannot guarantee across four cores, deletes the object, and checks both waiters came out with TX_DELETED. Two waiters rather than one so the loop takes its back edge as well as its body, and every wait is bounded in ticks so a suspension that never arrives fails the test instead of hanging it. threadx_trace_entry_update_test and threadx_thread_misaligned_stack_test are ports of the two tests that closed the equivalent gaps in common/src, and close fourteen more lines here: tx_block_allocate 123, 175, 182, 319 and 326, tx_byte_allocate 130, 210, 217, 359 and 366, tx_thread_system_suspend 504 and 560, tx_trace_object_register 221, and tx_thread_create 133. The one substantive change is core confinement. The trace test needs thread 0 to suspend and thread 1 to then release what it waits for; on four cores thread 1 gives the block back before thread 0 has suspended and the update block behind the suspension is never reached, so both threads are excluded from cores 1 to 3. The misaligned stack test needed no such change. threadx_byte_memory_long_search_test closes three of the eleven in tx_byte_pool_search. Lines 264, 267 and 270 are the TX_BYTE_POOL_MULTIPLE_BLOCK_SEARCH limit -- twenty on this port -- where a long search drops and retakes protection so that it cannot lock the other cores out for the whole walk. No byte pool in the suite ever had twenty fragments. This one is filled with small chunks until it refuses another and then has every second chunk released, so the free fragments are never adjacent and cannot be merged, and the request is larger than any of them but smaller than the pool's theoretical total, which is what makes _tx_byte_pool_search walk rather than refuse at the door. The layout is asserted rather than assumed: the test checks the fragment count and checks the probe request really does fail before the workers start, because either would otherwise turn it into a silent no-op. Eleven lines remain and they are not a to-do list. Eight are the delay loop in tx_byte_pool_search that fires when another thread claims the pool inside the window the search opens. The Linux SMP port serialises all four cores on one pthread mutex, so that window is an unlock immediately followed by a lock on that mutex, and glibc hands an uncontended mutex straight back to the thread that just released it: measured over 180,003 windows across three cores, with zero handovers. The shipped test therefore does 250 searches per worker rather than the sixty thousand that probe used, because the twenty-block threshold is crossed by the first search. The other three are in tx_thread_smp_utilities. Line 149 is a range guard placed after the shift it is meant to guard, so reaching it needs a shift by the width of the type; the fix is to move the check above the shift, matching the TX_MAX_PRIORITIES > 32 variant of the same function, and that belongs in its own change. Lines 1073 and 1074 need a mutex owner that is genuinely executing on another core when a waiter suspends, and three shapes were tried without producing one on this port. Measured twice before and twice after, every gcda deleted between runs and 570 of 570 tests passing each time: 5114 of 5178 both times before, 5167 of 5178 both times after. Branch coverage goes from 2768 to 2821 and 2823 of 3548. A floor of 99 needs 5127, so the ratchet lands with forty lines of headroom against a numerator that has been seen moving by two between runs. Assisted-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ff7fbe02f8 |
Added a GCC check for the Arm ports and ran it in CI (#675)
* Assembled the module ports, which no check had ever compiled
scripts/check_clang.sh globbed ports_module/*/gnu/src, which does not exist --
the module ports keep their assembly in module_manager/src. The [ -d ] guard
skipped it in silence, so 116 assembly files across nine Arm module ports were
assembled by no check, with either compiler, in the script whose own comments
state three times that "a port that is simply absent from the count reads as
covered". Stage 1 goes from 724 of 724 to 840 of 840; the feature-macro stage
had the same gap and goes from 412 files to 469.
Correcting the path exposed five defects, and only one of them was a build
failure. The other four assembled cleanly and did the wrong thing, because GAS
runs the C preprocessor on .S and not on .s:
ports_smp/cortex_a7_smp/gnu/src/tx_thread_smp_unprotect.s, the only .s in a
directory of twenty-one .S, ignored all four of its own feature macros. It
wrote the caller's LR into the protection structure on every unprotect -- a
store guarded by TX_MPCORE_DEBUG_ENABLE -- sent an unconditional SEV, and
returned through both BX lr and MOV pc, lr. Its cortex_a5_smp and
cortex_a9_smp siblings are .S.
ports_module/cortex_m33/.../tx_thread_stack_build.s emitted both arms of an
#ifdef TX_SINGLE_MODE_SECURE, so the non-secure LR value overwrote the secure
one and the secure build got the wrong frame.
ports_module/cortex_m23/.../tx_thread_context_{save,restore}.S carried the
POP {r0, lr} that check_clang.sh's own comment describes as the reason the
feature-macro stage exists. The 16-bit Thumb POP takes r0-r7 and pc only.
The identical fix already sits in ports/cortex_m23/gnu/src; the module copy
never got it because nothing scanned it.
ports_module/cortex_m23/.../tx_thread_secure_stack_initialize.S used MOV
rather than MOVS for an 8-bit immediate, latent behind TX_SINGLE_MODE_SECURE.
Both siblings in the same directory already use MOVS.
ports_module/cortex_a7/gnu/module_manager/src is the one that failed to
assemble, on GCC 14.3 as well as on LLVM: #define SYS_MODE was never
expanded, so #SYS_MODE reached the assembler as an undefined symbol.
Twenty-nine .s files under gnu trees are renamed to .S. Every one of them is
already named .S by the build scripts that compile it, so this repairs those
scripts rather than churning them -- ports_module/cortex_a7's build_threadx.bat
names all eighteen with a capital S, and works today only on a case-insensitive
filesystem. Renaming rather than converting the #defines to GNU assignments is
what fixes the #ifdef blocks as well as the constants; the assignments would
have fixed two files and left twenty-seven silently ignoring their macros.
Files with no preprocessor directives are left as .s: they are not broken, and
check_ports.sh gains a check that keeps them that way. Only the gnu trees are
checked there -- the IAR, Arm Compiler 5 and Keil assemblers preprocess .s
themselves, and about three hundred files in this repository rely on that.
Verified with both toolchains on the same tree: 840 of 840 assembled by
ATfE 22.1.0 and by arm-gnu-toolchain 14.3.rel1, all five stages of
check_clang.sh green, and check_ports.sh green including the reproducibility
check. The new check was shown to fail by planting a copy of the file it was
written for.
No regression test accompanies this. The assembly it covers is executed by no
host test, and the check itself going from 724 files to 840 is the coverage
AGENTS.md asks for -- together with the new check_ports.sh section, which is
what stops the class recurring.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Fixed the AArch64 samples, none of which had ever linked with GCC
Every AArch64 gnu example build failed at the sample link, all 27 of them --
13 under ports/ and 14 under ports_smp/:
libg.a(libc_a-init.o): in function `__libc_init_array':
undefined reference to `_init'
relocation truncated to fit: R_AARCH64_CALL26 against undefined
symbol `_init'
libg.a(libc_a-fini.o): in function `__libc_fini_array':
undefined reference to `_fini'
build_threadx_sample.sh links with -nostartfiles, which is correct for a port
carrying its own reset path, and that drops crti.o and crtn.o along with
everything else. startup.S calls __libc_init_array by design, and newlib's
implementation calls _init, which crti.o is what defines. The AArch32 scripts
are unaffected: they use nosys.specs and never reach __libc_init_array.
The fix links crti.o and crtn.o explicitly, bracketing the object list -- the
first must precede every .init contribution and the second must follow all of
them, so their position is load-bearing rather than stylistic. Both paths come
from the compiler's own -print-file-name, so nothing here hard-codes a
toolchain layout.
The atfe branch sets both to empty, deliberately: picolibc's __libc_init_array
does not call _init, those 27 images link today, and adding crti.o would change
a working link for no reason. That is also why check_clang.sh is green on these
and does not list them as expected to fail -- the LLVM path never reached the
gap, so nothing has ever linked them and failed.
Fixed in ports_arch/ARMv8-A/threadx/ports/gnu/example_build, which is the
single source for both the ports/ and ports_smp/ copies, then regenerated with
update.sh --port-sets tx,tx_smp. The 27 generated copies are in this commit
because ports_arch_check compares them.
Verified: all 27 link with arm-gnu-toolchain 14.3.rel1 aarch64-none-elf, where
0 of 27 did before; _init and _fini disassemble to the expected crti prologue
and crtn epilogue over a ret; check_clang.sh with ATfE 22.1.0 is still green on
all five stages, including the 42 script-driven example builds; check_ports.sh
is green including the reproducibility check.
No regression test: these are link-only example images that no host test
executes. What guards them is check_clang.sh's example stage today, and
check_gcc.sh's, which is the next change and is the reason this was found.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Added a GCC check for the Arm ports, which nothing had ever compiled
GCC is the project's declared default compiler (AGENTS.md, "The default
compiler for the project is GCC 14 on Linux"), it is what the gnu ports exist
for, and it is what nearly every downstream user builds with -- and nothing in
CI compiled a line of any port with it. The only cross-compilation check that
ran was the LLVM one, so the ATfE path was better guarded than the GNU one, on
ports whose directory is literally named gnu. ci_cortex_m covers four port
families; this covers forty.
Five stages, mirroring scripts/check_clang.sh stage for stage:
1. assemble every .S and .s of every Arm gnu port -- 840 files
2. assemble again the parts behind TX_ENABLE_VFP_SUPPORT,
TX_ENABLE_FIQ_SUPPORT, TX_LOW_POWER and
TX_ENABLE_EXECUTION_CHANGE_NOTIFY -- 469 files
3. compile common/src for one core per architecture profile -- 185 x 9
4. link the script-driven example builds -- 42
5. link the CMake-driven Cortex-R52 images -- 5
Two scripts rather than one with a --toolchain flag: the flag surface differs
(a prefixed driver against --target=), the C library differs, and the set of
examples that can link differs. Folding them together makes it easy to weaken
one check while working on the other.
Two toolchains, both required. Arm ships arm-none-eabi and aarch64-none-elf as
separate downloads and PORT_TARGET maps every port to one of exactly those two
triples, so --arm-none-eabi and --aarch64-none-elf each take a driver or the
directory holding it, defaulting to the environment and then to PATH. A missing
one is a hard error rather than a soft skip: letting a run cover half the tree
and still report "all checks passed" is the failure this script exists to end.
PORT_TARGET is copied verbatim from check_clang.sh, including its warning not
to prefix-match core names -- cortex_a5* also matches the AArch64 cortex_a53.
VFP_EXTRA is the one map that is not a copy, and check_clang.sh's comment about
it is false for GCC. That comment says the A-profile defaults are already
correct; arm-none-eabi-gcc defaults to -mfloat-abi=soft, which disables the FPU
outright, so every VFP file fails with "selected processor does not support
'vmrs r1,FPSCR' in ARM mode". -mfloat-abi=hard alone is the fix and is the
right one, because it selects the core's own default FPU rather than naming a
-d16 one -- which is the trap the clang script warns about, since the
A-profile paths save D16-D31. Cortex-R4 is the exception in both scripts and
for the same reason: its FPU is an option rather than part of the core, so an
explicit -mfpu is required. Every value was measured against 14.3.rel1.
Stage 4 *unsets* TOOLCHAIN rather than setting it. The example build scripts
already default to GNU, and a stray TOOLCHAIN=atfe from a developer's shell
would otherwise make this stage silently check the other compiler. It cleans
the example directories on both sides, because the success test is the
existence of sample_threadx.out rather than the driver's exit status, and a
stale image from a previous toolchain would report success. Failure logs are
printed unfiltered: a missing tool says "command not found", and GNU ld's
undefined-symbol lines carry no "error:" at all.
Every skip is printed by name with a reason, per the house rule check_clang.sh
states three times -- a port simply absent from the count reads as covered.
This script also says outright that arm9 and arm11 are Arm and are skipped for
having no PORT_TARGET entry, which the clang script's "not Arm" wording glosses.
Verified on this tree with arm-gnu-toolchain 14.3.rel1: all five stages green,
every count identical to check_clang.sh's on the same tree -- 840, 469, 185x9,
42, 5 -- in 4m28s.
The failure paths were tested, not assumed. A deliberately broken .S in a
module port is reported by name and line in stages 1 and 2 and exits 1, in
--quiet mode as well. Reverting the AArch64 _init/_fini fix on one port only
gives "FAIL: cortex_a53: example build produced no image", 41 of 42, and exit
1 -- and the log tail it prints contains no "error:" anywhere, which is why it
is not filtered. A missing or wrong toolchain path exits 1 naming which triple
was not found.
RISC-V is deliberately out of scope for this first version: both ports
assemble 8 of 8 with the project's own cmake flags, but adding them widens the
toolchain download and the review surface for a family that is not regressing.
No regression test accompanies this. The script is the test, it exercises no
runtime behaviour, and its own failure paths are exercised above.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Ran the GCC port check in CI, on dev as well as master
scripts/check_gcc.sh with nothing invoking it would be a script nobody runs.
This adds the workflow, modelled on clang_check.yml, and fixes a trigger gap in
that file at the same time.
One job, two cache steps. Arm ships AArch32 and AArch64 as separate downloads
and the script needs both, so two caches keep the checks list short and let a
single invocation see both compilers. The AArch32 cache path and key match
cortex_m's exactly, so the two workflows share one entry rather than each
holding its own copy of the same archive -- noted in a comment, because the
only symptom of breaking that is a slower run.
Both triggers name dev. A workflow that triggers only on master gates no pull
request anybody opens; that is the defect ports_arch_check.yml carries a
comment about, and it cost cortex_m three months of failing in seven seconds
unnoticed. push is included as well as pull_request so dev's own history has a
baseline and a bad squash-merge is caught rather than waiting for the next PR.
The checksum suffix is .sha256asc and it is not interchangeable with .sha256.
Arm publishes both for this release, and verified 26 Aug 2026, the .sha256 file
for arm-none-eabi contains a 32-character MD5 rather than a SHA-256, so
sha256sum -c on it fails with "no properly formatted checksum lines found".
.sha256asc is a plain sha256sum-format line for both triples. The plan warned
that this suffix had changed between releases; the sharper truth is that both
suffixes exist simultaneously and one of them is not a SHA-256 at all. Recorded
in a comment beside the step.
Verified before writing them in rather than copied: both archive URLs and both
checksum URLs resolve, the archives are xz, the checksum files are
sha256sum-format for .sha256asc, and the AArch64 archive extracts to
arm-gnu-toolchain-14.3.rel1-x86_64-aarch64-none-elf/bin/aarch64-none-elf-gcc,
which is the path the workflow builds.
The paths: lists are duplicated between push and pull_request rather than
shared through a YAML anchor, deliberately: GitHub Actions' parser does not
dependably honour anchors and the failure mode is the workflow refusing to
parse, which is the cortex_m failure again. Ten duplicated lines are cheaper.
clang_check.yml's paths: list was missing CMakeLists.txt, cmake/ and
common_smp/, so that check did not run when files it reads changed -- the
ports_smp example builds compile common_smp/src and its CMake stage reads the
toolchain file and the top-level project. Both lists are now identical apart
from each file's own name, and both say so.
cortex_m is kept rather than deleted, against the plan's recommendation. It
builds four ports *through CMake*, and that is the only thing exercising
cmake/cortex_m*.cmake and the top-level CMakeLists for the M profile; this
script's CMake stage covers cortex_r52 only. The overlap is the assembly and
the C sources, not the build system, so deleting it would lose coverage rather
than remove a duplicate. Said so in the workflow header.
The script is passed explicit toolchain paths rather than left to find the
drivers on PATH, so nothing about the runner image can decide which compiler
runs, and it prints both versions it resolved.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
5b94bad6a2 |
Fixed the AArch64 samples, none of which had ever linked with GCC (#673)
Every AArch64 gnu example build failed at the sample link, all 27 of them --
13 under ports/ and 14 under ports_smp/:
libg.a(libc_a-init.o): in function `__libc_init_array':
undefined reference to `_init'
relocation truncated to fit: R_AARCH64_CALL26 against undefined
symbol `_init'
libg.a(libc_a-fini.o): in function `__libc_fini_array':
undefined reference to `_fini'
build_threadx_sample.sh links with -nostartfiles, which is correct for a port
carrying its own reset path, and that drops crti.o and crtn.o along with
everything else. startup.S calls __libc_init_array by design, and newlib's
implementation calls _init, which crti.o is what defines. The AArch32 scripts
are unaffected: they use nosys.specs and never reach __libc_init_array.
The fix links crti.o and crtn.o explicitly, bracketing the object list -- the
first must precede every .init contribution and the second must follow all of
them, so their position is load-bearing rather than stylistic. Both paths come
from the compiler's own -print-file-name, so nothing here hard-codes a
toolchain layout.
The atfe branch sets both to empty, deliberately: picolibc's __libc_init_array
does not call _init, those 27 images link today, and adding crti.o would change
a working link for no reason. That is also why check_clang.sh is green on these
and does not list them as expected to fail -- the LLVM path never reached the
gap, so nothing has ever linked them and failed.
Fixed in ports_arch/ARMv8-A/threadx/ports/gnu/example_build, which is the
single source for both the ports/ and ports_smp/ copies, then regenerated with
update.sh --port-sets tx,tx_smp. The 27 generated copies are in this commit
because ports_arch_check compares them.
Verified: all 27 link with arm-gnu-toolchain 14.3.rel1 aarch64-none-elf, where
0 of 27 did before; _init and _fini disassemble to the expected crti prologue
and crtn epilogue over a ret; check_clang.sh with ATfE 22.1.0 is still green on
all five stages, including the 42 script-driven example builds; check_ports.sh
is green including the reproducibility check.
No regression test: these are link-only example images that no host test
executes. What guards them is check_clang.sh's example stage today, and
check_gcc.sh's, which is the next change and is the reason this was found.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
9c32abb17d |
Assembled the module ports, which no check had ever compiled (#672)
scripts/check_clang.sh globbed ports_module/*/gnu/src, which does not exist --
the module ports keep their assembly in module_manager/src. The [ -d ] guard
skipped it in silence, so 116 assembly files across nine Arm module ports were
assembled by no check, with either compiler, in the script whose own comments
state three times that "a port that is simply absent from the count reads as
covered". Stage 1 goes from 724 of 724 to 840 of 840; the feature-macro stage
had the same gap and goes from 412 files to 469.
Correcting the path exposed five defects, and only one of them was a build
failure. The other four assembled cleanly and did the wrong thing, because GAS
runs the C preprocessor on .S and not on .s:
ports_smp/cortex_a7_smp/gnu/src/tx_thread_smp_unprotect.s, the only .s in a
directory of twenty-one .S, ignored all four of its own feature macros. It
wrote the caller's LR into the protection structure on every unprotect -- a
store guarded by TX_MPCORE_DEBUG_ENABLE -- sent an unconditional SEV, and
returned through both BX lr and MOV pc, lr. Its cortex_a5_smp and
cortex_a9_smp siblings are .S.
ports_module/cortex_m33/.../tx_thread_stack_build.s emitted both arms of an
#ifdef TX_SINGLE_MODE_SECURE, so the non-secure LR value overwrote the secure
one and the secure build got the wrong frame.
ports_module/cortex_m23/.../tx_thread_context_{save,restore}.S carried the
POP {r0, lr} that check_clang.sh's own comment describes as the reason the
feature-macro stage exists. The 16-bit Thumb POP takes r0-r7 and pc only.
The identical fix already sits in ports/cortex_m23/gnu/src; the module copy
never got it because nothing scanned it.
ports_module/cortex_m23/.../tx_thread_secure_stack_initialize.S used MOV
rather than MOVS for an 8-bit immediate, latent behind TX_SINGLE_MODE_SECURE.
Both siblings in the same directory already use MOVS.
ports_module/cortex_a7/gnu/module_manager/src is the one that failed to
assemble, on GCC 14.3 as well as on LLVM: #define SYS_MODE was never
expanded, so #SYS_MODE reached the assembler as an undefined symbol.
Twenty-nine .s files under gnu trees are renamed to .S. Every one of them is
already named .S by the build scripts that compile it, so this repairs those
scripts rather than churning them -- ports_module/cortex_a7's build_threadx.bat
names all eighteen with a capital S, and works today only on a case-insensitive
filesystem. Renaming rather than converting the #defines to GNU assignments is
what fixes the #ifdef blocks as well as the constants; the assignments would
have fixed two files and left twenty-seven silently ignoring their macros.
Files with no preprocessor directives are left as .s: they are not broken, and
check_ports.sh gains a check that keeps them that way. Only the gnu trees are
checked there -- the IAR, Arm Compiler 5 and Keil assemblers preprocess .s
themselves, and about three hundred files in this repository rely on that.
Verified with both toolchains on the same tree: 840 of 840 assembled by
ATfE 22.1.0 and by arm-gnu-toolchain 14.3.rel1, all five stages of
check_clang.sh green, and check_ports.sh green including the reproducibility
check. The new check was shown to fail by planting a copy of the file it was
written for.
No regression test accompanies this. The assembly it covers is executed by no
host test, and the check itself going from 724 files to 840 is the coverage
AGENTS.md asks for -- together with the new check_ports.sh section, which is
what stops the class recurring.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
147754cc86 |
Enforced a coverage floor on the merged report (#667)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The coverage summary reported a percentage and could not fail. Coverage could fall from 99.97% to anything at all and every check stayed green, against an AGENTS.md that asks for 100% test coverage -- a stated requirement measured with a gauge that had no failure mode. CodeCoverageSummary already takes thresholds and fail_below_min; neither was set. Both are now, through a new coverage_thresholds input on the template, because the two suites do not sit at the same figure: ThreadX 99, SMP 98. Three things were probed against the pinned action on a runner before picking those numbers, using the real merged.xml files from the dev push run of #666. The floor compares the line rate and nothing else. That mattered because branch coverage is around 78% in both suites while line coverage is 98.8-100%, so a floor aimed at the line figure would have been an immediate red wall had it tested branches or the lower of the two. The ThreadX report at 100.00% lines and 77.67% branches clears a floor of 99. The thresholds are whole numbers. '99.9 100' -- the value this was meant to be -- is rejected with 'System.ArgumentException - Threshold parameter set incorrectly.', and the step fails whether or not fail_below_min is set. So the choice is 99 or 100 with nothing between. 100 would fail on a race. tx_thread_system_resume.c:529 is reached by timing rather than by construction and flaps between runs of the same green tree, which is why #666 left it; 4502/4503 fails a floor of 100 and clears one of 99. A coverage gate that goes red on a coin toss is how coverage gates get switched off. SMP is 5114/5178 lines, 98.76%, with 64 uncovered lines across 11 files of common_smp/src -- #666 closed the equivalent gaps in common/src only. A shared floor of 99 would have failed that job on every run while ThreadX passed. One limit is recorded in the file rather than fixed: an empty report reads as 100%. gcovr writes line-rate="1.0" beside lines-valid="0" when it finds no data, and the action prints 'Line Rate = 100% (0 / 0)' and passes any floor. The check for that is the emptiness assertion #664 put in each suite's coverage.sh, not this one. Also corrected two stale filenames in the deploy job's comment: since #665 each coverage artifact carries merged.xml, not default_build_coverage.xml. Verified on the runner. Assisted-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f89d65f041 |
Covered the trace entry update paths and the misaligned stack adjustment (#666)
Against the merged coverage report -- every build configuration instrumented and
unioned -- eighteen lines of common/src were uncovered. Seventeen of them were
one missing scenario rather than eighteen separate gaps.
tx_block_allocate, tx_byte_allocate, tx_thread_system_suspend and
tx_thread_system_resume each carry blocks under TX_ENABLE_EVENT_TRACE that go
back and patch a trace entry once the call has done its work, all of the shape:
if (entry_ptr != TX_NULL)
{
if (time_stamp == entry_ptr -> tx_trace_buffer_entry_time_stamp)
entry_ptr comes from _tx_trace_buffer_current_ptr, which stays TX_NULL until
tx_trace_enable is called at run time. Building with TX_ENABLE_EVENT_TRACE is
not enough, and exactly one test in the suite enables tracing --
threadx_trace_basic_test -- which tests the enable API itself and never calls
either allocator. So those blocks sat in the report's denominator and never in
its covered set.
threadx_trace_entry_update_test enables tracing and then drives both allocators
twice each, once on the path that succeeds immediately and once through a
suspension that a second thread satisfies, since each allocator carries one
update block on either side. It then sleeps so that the last runnable thread
suspends with nothing ready to take over: tx_thread_system_suspend lines 345 and
351 are on the branch that sets _tx_thread_execute_ptr to TX_NULL, and the two
allocator suspensions never reach it because the other thread was always ready.
The same test closes tx_trace_object_register's TX_NULL name break by creating a
semaphore with no name. A semaphore and not a thread deliberately: for
TX_TRACE_OBJECT_TYPE_THREAD the register function dereferences the pointer it is
given to read the thread's priority, so that type needs a real TX_THREAD behind
it. threadx_trace_basic_test makes the equivalent call only under
ifndef TX_ENABLE_EVENT_TRACE, against the no-op stub.
threadx_thread_misaligned_stack_test covers the remaining line,
tx_thread_create.c:136, where a stack that does not begin on a ULONG boundary
costs a ULONG of size so that rounding the start up cannot run past the end of
the caller's buffer. Every other test hands tx_thread_create an aligned stack.
That line is compiled only under TX_ENABLE_STACK_CHECKING, so it is absent from
three of the five configurations' reports rather than uncovered in them, and it
was verified under stack_checking_build.
Measured on the merged report, all 480 tests passing and 5 of 5 configurations
green: 4485 of 4503 covered before, 4502 of 4503 after, denominator unchanged.
One line remains, tx_thread_system_resume.c:529, and it is the report's last
flapping line rather than a standing gap -- two clean runs of the same tree gave
4503 of 4503 and 4502 of 4503. Reaching it by construction was tried twice and
failed both times, so it is left alone here. It needs the preempt disable flag
and the system state both clear, and tx_thread_resume raises the preempt disable
flag before calling _tx_thread_system_resume, as do the put and send paths;
creating a higher priority auto-start thread from thread context does not raise
it but does not reach the check either, which a probe showed is executed only
during initialisation, with the system state at TX_INITIALIZE_IN_PROGRESS.
Assisted-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
3e85bbd431 |
Instrumented every build configuration and merged their coverage (#665)
Only default_build_coverage carried -fprofile-arcs, because the gate was the build type and it is the only one of five whose name ends in _coverage. The other four build and run all their tests and their coverage was discarded. That is not redundancy thrown away: each configuration selects a different set of TX_ feature macros, so the code the other four compile is absent from the denominator rather than uncovered in it. TX_COVERAGE instruments a build regardless of its name, defaulting to OFF so a single configuration built by hand behaves as before. coverage.sh gains a --merge mode that unions the per-configuration JSON tracefiles, and cmake_bootstrap.sh runs it after the test loop so a local run produces the same merged report CI reads. The template sets TX_COVERAGE for build and test, and coverage_name moves to the merged report. Measured on the ThreadX suite, all 480 tests passing: default_build_coverage 3827 valid 3827 covered disable_notify_callbacks_build 3767 3766 stack_checking_build 3857 3856 stack_checking_rand_fill_build 3862 3861 trace_build 4123 4108 merged 4503 4487 The denominator grows by 676 lines, 17.7%, and the figure moves from 99.97% to 99.64%. The second one is honest, and the drop is the point rather than a regression: the denominator now includes code the old report never counted. The union also contains a file the old report did not contain at all -- tx_thread_stack_error_handler.c compiles only under TX_ENABLE_STACK_CHECKING, so it was not listed at 0%, it was simply absent. 177 files becomes 178. Coverage collection moved out of test() and now runs after the test loop, one configuration at a time. gcov writes its intermediate gcov files into the directory gcovr is rooted at, and coverage.sh roots every configuration at the repository root so filenames come out repo-relative. Five concurrent gcovr processes therefore share one scratch directory and delete each other's output: the first full run of this change passed all 480 tests and produced no report for three of the five configurations. Measured both ways -- two gcovr rooted at the repository root fail concurrently and succeed in sequence. CI would not have caught it, because test_tx.sh sets CTEST_PARALLEL_LEVEL=1 and takes the serial branch. Per-configuration output moved under coverage_report/per_configuration/ and is excluded from the Pages artifact. The deploy job merges the ThreadX and SMP artifacts into one tree and every configuration directory has the same name in both, so left at the top level one suite's would overwrite the other's on the published site. On the SMP suite, an earlier run of this change saw trace_build fail threadx_smp_time_slice_test and then hang, which raised the question of whether -fprofile-arcs perturbs a timing-sensitive test. It does not. Sixteen runs settle it, and the decisive one is that threadx_smp_time_slice_test failed ERROR #31 -- twice in a row under --repeat until-pass:2 -- on an uninstrumented build, in the exact shape CI runs, while three instrumented runs of that shape passed 5 of 5. In the CI shape, CTEST_PARALLEL_LEVEL=1 run.sh test all: TX_COVERAGE=OFF 3 runs 2 green, one ERROR #31 310 s TX_COVERAGE=ON 3 runs 3 green, 5/5 each 325-329 s So the test is a pre-existing flake on dev and instrumenting all five costs about 5% of the suite's wall clock. Separately, and also in both instrumented and uninstrumented builds, run.sh's parallel branch -- what a developer gets typing run.sh test all with no CTEST_PARALLEL_LEVEL -- hangs under its own load, four times in twelve runs. Several SMP tests create 1024 ThreadX threads by construction and the Linux port backs each with a pthread, so five configurations at once put on the order of 5000 threads on the machine. CI sets CTEST_PARALLEL_LEVEL=1 and does not take that branch. Assisted-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
b756220c43 |
Fixed the coverage report's paths and scoping (#664)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The Cobertura XML embedded absolute machine paths, and the flag that looked like it scoped the report to one build configuration was doing nothing at all. Both coverage.sh scripts had the defect; both are fixed here, because the SMP report is published to the same Pages site as the ThreadX one. Paths. -r was the build directory and -f pointed outside it, so gcovr could not express the sources relative to the root and fell back to absolute paths. The result named files as /home/runner/work/threadx/threadx/common/src/... while the <source> element beside them said build/default_build_coverage, so the two halves of the same file disagreed and nothing could map coverage back to the repository. -r is the repository root now and -f an absolute path beneath it, which gives filename="common/src/tx_block_allocate.c". Both must be absolute: -r ../../.. -f common/src produces a report of zero files and exits 0, which is the worst failure mode available here. Scoping. --object-directory does not restrict which gcda files are found -- it tells gcovr how to get from a gcda file back to the compiler's working directory. Pointed at an empty directory it still produced the full 177-file report. That was harmless only by accident, because -r build/$1 constrained the search instead; moving -r to the repository root removes that accident, so the two changes have to land together. Measured, with a second instrumented configuration deliberately made sparser than the first: scoped by the positional search path 3221 of 3827 lines -- the truth no search path, -r at the repo root 3827 of 3827 -- silently merged --object-directory at the sparse tree 3827 of 3827 -- scopes nothing So the search path is load-bearing, and it matters ahead of instrumenting all five configurations: without it each configuration would have reported the union as its own. An empty report is not an error to gcovr -- it warns and exits 0 -- and it carries line-rate="1.0" next to lines-valid="0", so a consumer reads no data at all as fully covered. No coverage threshold can catch that, since an empty report passes any threshold. Hence the explicit assertion that the report has content, which fires with exit 1 on an object directory that exists but is empty, where the old shape returned 177 files and exit 0. Also says out loud that ports/linux/gnu/src is deliberately outside the filter. gcno files exist for it and it is dropped without a word today. Number-neutral, and that was the test. Over the same frozen gcda, changing only the gcovr invocation: ThreadX 3827 of 3827 lines and 1993 of 1994 branches across 177 files, SMP 4739 of 4791 and 2417 of 2430 across 185, before and after alike. Same answers on gcovr 7.0, 8.3 and 8.6, so the change is not wedged to the current pin. End to end through run.sh, 96 of 96 and 110 of 110 pass with the reports written. Assisted-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
2d9b9f7417 |
Bumped gcovr off the 4.1 pin it had been held on since 2018 (#663)
The coverage tooling was pinned to gcovr 4.1, released in 2018, and that version is missing the two options the coverage work needs next: --json and --add-tracefile, which is how the five build configurations get merged into one report. This moves the pin to 8.6, the current release. The pin stays exact, and it stays hand-moved: it lives in a shell script, and no Dependabot ecosystem can parse that. Isolated deliberately, so that a movement in the coverage number caused by the tool could not be confused with one caused by a later change. Measured on the default_build_coverage tree of test/tx, over the same gcda with the same gcov, varying only the gcovr version: gcovr lines-valid branches-valid files 4.1 3827 1994 177 7.0 3827 1994 177 8.3 3827 1994 177 8.6 3827 1994 177 So the denominator does not move with the tool at all, and this bump moves no number. The plan this came from expected 3822 to become 3827; that figure does not reproduce, under gcc-13 or gcc-14, with or without --object-directory. The only variant that changes the count is dropping the -f filter, which collapses the report to nothing. Two things found while measuring, both recorded because they matter to what comes next. The coverage numerator is not deterministic. On an identical tree with an identical compiler, three consecutive runs of the full suite -- all 96 tests passing every time -- reported 3826, 3827 and 3827 covered lines. The line that flickers is tx_thread_system_resume.c:529, the preemption path of _tx_thread_system_resume, and it takes its guarding branch with it. It has been described as never executed; it is executed on some runs and not others. A coverage floor has to be set with that in mind, and the honest fix is a test that takes the path deliberately. Reading gcc-13 output, the compiler the runners actually use, gcovr 8.6 runs the existing coverage.sh unchanged: Cobertura XML and 181 HTML files, same 177 classes. --xml-pretty and --object-directory still work on 8.6 but are now deprecated aliases for --cobertura-pretty and --gcov-object-directory, worth knowing for whoever removes --object-directory next. Assisted-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
b6a00a2014 |
Added the Dependabot configuration the pinned actions need (#662)
The action references were pinned to commit SHAs in #660, and a SHA pin with nothing moving it is worse than a floating tag -- it holds CI on whatever was current the day it was written. That is exactly how actions/cache@v1 stayed in ci_cortex_m.yml until GitHub began auto-failing every request that used it. The drift measured before that catch-up: download-artifact four majors behind, checkout and upload-artifact three each, cache and upload-pages-artifact two, with nothing ever reporting it. This closes the loop, and the reference to .github/dependabot.yml that #660 left in each workflow's pinning comment. Weekly, github-actions only. Patch and minor are grouped into one pull request because they are the routine traffic and a queue reviewed one item at a time is a queue that gets ignored. Majors stay ungrouped, one each, because every breaking change this repository has met in an action has been a major. Two choices worth stating rather than leaving to be rediscovered. target-branch is dev. Dependabot reads this file from the default branch, which is master, but master is deliberately kept behind dev and pull requests belong where the regression suites gate them. The consequence is that landing this on dev arms it without firing it: nothing happens until a release merge carries the file to master. Setting target-branch also opts out of Dependabot security updates, which only run against the default branch -- a small cost for this ecosystem, since an action advisory arrives as an ordinary bump on the weekly run, but a real one. The pull-request limit is raised from the default five to ten. Nine actions are in use, and five would hold majors back with nothing saying that it had. No other ecosystem is configured, deliberately: external dependencies are forbidden, there are no submodules, and the one pinned tool -- gcovr in scripts/install.sh -- lives in a shell script no ecosystem can parse, so that pin keeps moving by hand. No sibling eclipse-threadx repository has a Dependabot configuration, so this sets the pattern rather than following one. The dependencies label it uses already exists here. Assisted-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
3d852eb451 |
Pinned every action to a commit SHA, and moved them off Node 20 (#660)
Node 20 is removed from the GitHub runners on 16 September 2026. Every run
in this repository currently emits the deprecation warning for it, naming
actions/checkout, actions/configure-pages, actions/upload-artifact,
LouisBrunner/checks-action and marocchino/sticky-pull-request-comment among
others. After that date those actions stop working rather than warning, so
this is a deadline and not housekeeping.
Every action is now referenced by a 40-character commit SHA with the version
in a trailing comment. A tag can be repointed at any commit; a SHA cannot, so
this is what makes "which code ran in our CI" answerable from the repository
rather than from whatever the tag meant at the time. The versions were behind
by as much as four majors -- download-artifact was on v4.3.0 against v8.0.1 --
because nothing in this repository has ever reported that an action moved.
Compatibility was checked against each new action.yml rather than assumed,
for every input this repository actually passes:
checkout submodules is unchanged
cache path and key are unchanged
upload-artifact name, path and retention-days are unchanged
download-artifact pattern, merge-multiple and path are unchanged
configure-pages takes no input here, and none became required
deploy-pages still exposes page_url, which the job reads
upload-pages-art. path is unchanged
checks-action token, name, conclusion, output and
output_text_description_file all survive v2 to v3
sticky-comment header and path survive v2 to v3, and the new
GITHUB_TOKEN input defaults to github.token, which is
what v2 used implicitly
delete-artifact name survives v5 to v6, and useGlob still defaults to
true, so the coverage_report-* glob from #655 still
matches
CodeCoverageSummary already current at v1.3.0; pinned, not moved
The artifact pair moves together, as it must. The round trip was verified on
a runner before this commit: upload-artifact v7 to download-artifact v8,
through the pattern and merge-multiple selection #655 introduced, filtered 4
artifacts to 2 and produced exactly the tree the deploy expects.
Two behaviour changes worth knowing. download-artifact v8 adds a
digest-mismatch input defaulting to error, so a corrupted artifact now fails
the job instead of passing through -- the right default, but a change.
upload-artifact v6 and above require a runner of at least 2.327.1, which the
hosted runners satisfy and a self-hosted runner would need checking for.
Assisted-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
042049a7b2 |
Revived the Cortex-M build, which had compiled nothing since June (#653)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
Two defects, and the second hid the first. This workflow triggered on master only, for both push and pull_request, while dev is the integration branch. So it gated no pull request that anybody opened. That is the same defect ports_arch_check.yml carries a comment about, where it cost eight months of ports drifting from ports_arch unnoticed, and regression_test.yml has it too. And it had not compiled anything since at least 2026-06-08. Every run since then failed in six to eight seconds at "Prepare all required actions", before checkout, because GitHub automatically fails any request that uses actions/cache@v1. The last run of any kind was 2026-06-30. A workflow that fails in seven seconds is normally noticed within the day; this one was not, because of the first defect. The two together meant the project's only job that cross-compiles a port with GCC had been reporting nothing at all. The toolchain now follows clang_check.yml rather than third party actions: a pinned release fetched directly from Arm, verified against the published sha256asc, and cached with actions/cache@v4. That also moves the compiler off 9-2019-q4, a 2019 release, onto a version matching the GCC 14 default this project states. The ninja install is guarded on ninja being absent rather than run unconditionally, because scripts/install.sh already carries a long comment about apt-get update stalling for over two hours and taking a whole regression run with it. fail-fast is off so that one port failing still reports the other three. Verified before committing, with the pinned 14.3.rel1 toolchain: the download and sha256sum -c sequence in the install step was run as written, the archive extracts to the directory the PATH step expects, and all four ports configure and build clean with zero warnings. This covers four ports of the forty under ports/ that have a gnu directory. Widening it to every Arm gnu port is separate work. Assisted-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
7959aef3bc |
Kept the coverage report from the runs that most need one (#659)
A failing test threw away coverage that had already been collected, and the run whose behaviour changed is exactly the run whose coverage is worth reading. Measured on the failing run of 2026-08-18: it uploaded test_reports for all three suites and no coverage_report artifact at all. Two causes, and the workflow one is the smaller of them. cmake_bootstrap.sh runs under set -e, so a failing ctest aborted test() before ./coverage.sh was reached. The gcda files exist by that point, so nothing was missing except the step that reads them. ctest's status is now captured and returned at the end, and the summary grep is allowed to fail rather than being the thing that stops the coverage behind it. The serial branch of the test dispatch collected no status either, so under set -e the first failing configuration stopped the remaining four from being tested at all -- and their coverage from being collected. That was cheap while the suites ran in parallel, because the parallel branch already collects exit codes from its background jobs. Moving to serial execution in #643 quietly made one failure cost the other four configurations. The serial branch now collects status the same way the parallel branch does. With those fixed the report exists, so the workflow steps that publish it no longer skip on failure. They are guarded with !cancelled() rather than always(), so a cancelled run still stops promptly, which is the idiom deploy_code_coverage already uses. The ${{ }} wrapping is required and not decoration: a bare ! opens a YAML tag, and the file will not parse without it. Verified locally by replacing one test binary with a stub that exits 1: before the failing run of 2026-08-18 produced no coverage_report artifact after run.sh test default_build_coverage exits 8, and produces coverage_report/default_build_coverage.xml with 177 files and 3804 of 3827 lines after run.sh test all exits 8, and all five configurations run rather than stopping at the first The failure still fails. Only the reporting around it changed. Assisted-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
adc6469b91 |
Published the coverage report instead of everything the run produced (#655)
The download step in deploy_code_coverage asked for the artifact named
${{ steps.artifact.outputs.coverage_report }}. That output is set by the
"Coverage Report name" step of run_tests, which is a different job, and the
steps context does not cross jobs. So the expression evaluated to the empty
string and the action took its documented path for an unspecified name:
No input name, artifact-ids or pattern filtered specified,
downloading all artifacts
Total of 4 artifact(s) downloaded
The four are the two coverage reports and the two test_reports bundles of
JUnit XML, each extracted into a directory named after the artifact. The
next step uploads the lot to Pages, so the published site has carried the
test reports alongside the coverage, one directory deeper than intended,
under a path containing a run timestamp that changed on every publish. Any
link to a coverage report broke the next time one was published.
Selecting by pattern with merge-multiple fixes both halves: the pattern
excludes the test_reports bundles, and merging puts the contents of the two
coverage artifacts directly into coverage_report rather than under a
directory named for each. Each artifact holds one directory named for its
suite, renamed from default_build_coverage by "Prepare Coverage GitHub
Pages", so the result is the two suite directories the deploy expects and
the timestamped artifact name no longer appears in the published path.
Verified on a runner rather than reasoned about, with an isolated workflow
that uploads artifacts shaped like the real ones and downloads them both
ways:
OLD coverage_report/coverage_report-<epoch>-ThreadX/ThreadX/index.html
coverage_report/coverage_report-<epoch>-ThreadX/default_build_coverage.xml
coverage_report/coverage_report-<epoch>-SMP/SMP/index.html
coverage_report/coverage_report-<epoch>-SMP/default_build_coverage.xml
coverage_report/test_reports SMP/results.xml
coverage_report/test_reports ThreadX/results.xml
NEW coverage_report/ThreadX/index.html
coverage_report/SMP/index.html
coverage_report/default_build_coverage.xml
Both artifacts carry a default_build_coverage.xml and the merge means one
overwrites the other, which the run above also shows. That file is consumed
by CodeCoverageSummary back in run_tests and is not read here, so it is
untidy rather than wrong, and it is called out in a comment.
The delete step is fixed in the same place and for a related reason. The
artifacts are named coverage_report-<epoch>, useGlob defaults to true in
this action, and as a glob "coverage_report" matches only the literal
string. It has been deleting nothing, without failing, and retention-days: 1
on the upload is what has actually been clearing these up.
Assisted-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
eabdb86409 |
Matched gcov to the compiler that produced the coverage data (#658)
gcov reads a data format tied to the compiler that produced it. coverage.sh took whatever gcov was first on PATH, which was fine while the compiler was also whatever was first on PATH. #656 made cmake/linux.cmake honour CC and #657 made a compiler switch actually reconfigure the build, so that assumption no longer holds, and the first person to use the new capability would have hit this. Measured on dev with both of those merged: CC=gcc-14 ./run.sh build default_build_coverage # succeeds ./coverage.sh default_build_coverage # exit 64 gcov says why, if asked directly: tx_block_allocate.c.gcno:version 'B42*', prefer 'B33*' gcovr turns that into "GCOV returncode was 3" and exits 64 through a Python traceback, after the tests have already passed. It reads like a coverage bug rather than a toolchain mismatch, which is the part that would have cost someone an afternoon. gcov is now derived from CC rather than found on PATH, so the caller sets one variable instead of remembering two. GCOV still overrides, for a toolchain that does not follow the gcc/gcov naming, and a derived gcov that does not exist is reported as such instead of surfacing as a traceback. Verified, tx and smp, before and after: CC=gcc-14 was exit 64, now exit 0, 177 files and 1527/3827 lines CC unset exit 0, 177 files and 1527/3827 lines, unchanged CC=gcc-99 exit 1 naming gcov-99 and CC, rather than a traceback GCOV=gcov-14 with CC=gcc-99, exit 0, so the override still wins A mismatched pairing still fails, deliberately: reading a gcc-14 tree with the default gcc-13 gcov is exit 64 as before. Producing a number from mismatched data would be worse than refusing. Assisted-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
a4397e133a |
Reconfigured the build when the requested compiler changes (#657)
CMake records the compiler it detected inside the build directory and keeps using it on every later configure. Since cmake/linux.cmake began honouring CC, that made a compiler switch silently ineffective: CC=gcc-14 ./run.sh build <cfg> against an existing build directory printed "ninja: no work to do", exited 0, and left the previous compiler in place. Anyone verifying a change against a second compiler would have been reading stale results while being told the build had succeeded. generate() now compares the compiler recorded in the build directory with the one currently requested, and reconfigures from scratch when they differ. The comparison uses the path as CMake records it, which is the unresolved path as given, so /usr/bin/gcc matches command -v gcc rather than the versioned target its symlink points at. build_libs() gets the same treatment. Only the C compiler is consulted: both trees that use this script declare LANGUAGES C, so no CXX compiler is ever detected. Nothing is reconfigured unless the compiler actually changed, so repeat builds stay incremental and the default path is unchanged. Assisted-by: Claude Code (Opus 5) |
||
|
|
5a68c9da4b |
Allowed the Linux toolchain file to accept a compiler override (#656)
cmake/linux.cmake set CMAKE_C_COMPILER and CMAKE_CXX_COMPILER unconditionally. CMake reads a toolchain file before it consults CC and CXX, and a plain set() in a toolchain file also takes precedence over -DCMAKE_C_COMPILER, so neither of the two usual ways to pick a compiler had any effect: the tree could only ever be built with whatever /usr/bin/gcc happened to point at. That matters because AGENTS.md names GCC 14 as the project's default compiler on Linux, while distributions still ship an older gcc as the default for some time. Selecting GCC 14 previously meant either editing this file or changing the machine's system-wide default. Both variables now fall back to gcc and g++ only when nothing else has been specified, so the default build is byte-for-byte what it was, while -DCMAKE_C_COMPILER=gcc-14 or CC=gcc-14 now work as expected. The binutils variables in this file are left alone: they are unused on the Linux target, so guarding them would be unrelated churn. Assisted-by: Claude Code (Opus 5) |
||
|
|
9218bad4bc |
Kept the coverage publish on master, where the environment allows it (#654)
Running the regression suites on dev (#652) was meant to test the branch the pull requests target. It changed what gets published as well, which was not intended and does not work: the first push to dev after that merge failed with Branch "dev" is not allowed to deploy to github-pages due to environment protection rules. All three suites passed in that run -- tx, smp and freertos. The only failure was deploy / deploy_code_coverage, rejected before it ran, because the github-pages environment restricts deployments to master. The guard goes here rather than in the environment settings, because the environment rule is doing its job. Which branch the published coverage report describes is a deliberate decision, and moving it from master to dev is a change worth making on purpose rather than as a side effect of a trigger fix. Doing so needs the environment setting relaxed as well as this line removed. The per-suite deploy_code_coverage jobs need no guard: tx, smp and freertos all pass skip_deploy: true, and regression_template.yml already restricts that job to push and workflow_dispatch. Only the deploy job, which is the one that publishes, was reaching the environment. Assisted-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
977e14e776 |
Ran the regression suites on dev, where the pull requests actually are (#652)
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The ThreadX, SMP and FreeRTOS-compatibility suites triggered on master only, for both push and pull_request. dev is the integration branch, so these suites gated no pull request that anybody opened: the last dev run of any kind was a manual workflow_dispatch on 2026-08-18. This is the same defect ports_arch_check.yml already carries a comment about, where it cost eight months of ports drifting from ports_arch unnoticed. ci_cortex_m.yml has it too and is handled separately. That 2026-08-18 run failed, which is the reason to check before switching this on rather than after. Two tests failed: threadx_timer_simple_test in the ThreadX suite, with ERROR #28, and threadx_thread_priority_change in the SMP suite, with a timeout. Both were fixed two days later -- the first by running the suites one test at a time (#643), which is what a timer test failing only under parallel load wants, and the second by #647 by name. Verified before this commit rather than assumed: both suites were re-run on this tree, and all 1030 tests pass across all ten build configurations, the ThreadX suite in 34 to 64 seconds per configuration and the SMP suite in 61 to 63. The 2026-08-18 run took 36m19s, of which a single test that has since been given a budget accounted for 439 seconds. No paths filter is added deliberately. The suites build the linux port, so a filter would have to enumerate what cannot affect them, and the failure mode of getting that list wrong is a regression that merges because the filter excluded the file that caused it. The deploy job needs no guard: regression_template.yml already restricts deploy_code_coverage to push and workflow_dispatch, and restricts the coverage PR comment to pull requests from the repository itself, so neither fires for a pull request from a fork. Assisted-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
e46b1b0787 |
Asked the wait abort test for three windows, and failed a run that reached none (#649)
Four CI runs of the same tree, twenty configuration-runs in total, show this
test's budget being reached far more often than the first green run suggested,
and a pass being reported every time it was:
trace_build 3 of 10 windows in 121 seconds
disable_notify 3 of 10 windows in 121 seconds
default_coverage 4 of 10 windows in 121 seconds
stack_checking 7 of 10 windows in 121 seconds
trace_build 0 of 10 windows in 121 seconds
disable_notify 7 of 10 windows in 121 seconds
stack_checking 3 of 10 windows in 121 seconds
Seven of twenty, and the shortfall message only ever reaches an artifact:
ctest is run with --output-on-failure, so a passing test's output is not in the
job log at all. The suite has been quietly losing most of this test's coverage
in whole configurations and reporting green.
The loop runs in two modes, not one. A window arrives in milliseconds in the
fast mode, and costs between 17 and 40 seconds in the slow one, with nothing in
between across those twenty runs. Ten windows are therefore unreachable inside
any budget worth having: at 40 seconds each that is 400 seconds, and the
unbounded runs measured before any of this took up to 726. Raising the budget
to cover the slow mode would trade a quiet loss of coverage for five
configurations approaching the sixty minute step timeout.
So ask for what a run can reach. Three windows cost 51 to 120 seconds in the
slow mode and under a second in the fast one, and the later hits repeat what
the first ones establish, so what is given up is small. The budget goes to 180
seconds because three windows at the worst rate measured is exactly the 120 it
was, which would have truncated at two.
The count is printed on every run rather than only on a short one. A number
that appears only on shortfall cannot be told apart from a number nobody
recorded.
Reaching the window no times at all is a different matter, and was the worst of
the seven. The check after the loop compares semaphore bookkeeping that a
window has to have touched to mean anything, so a run that reached none of them
compares a counter against the value it was initialised to and reports a pass
having verified nothing. That run now keeps trying to a 300 second ceiling, and
fails if it still has not reached the window. A genuine resonance that holds
for five minutes is worth a failure; the old behaviour was worth nothing.
The SMP copy keeps its count of twenty. It reaches them in under half a second
in all five of its configurations, in all four runs, so the slow mode has never
been observed there and the coverage is free. Both copies get the ceiling and
the unconditional report, so the logic stays identical between them.
Verified locally on all five configurations: the test reaches 3 of 3 in 5 to 14
seconds, and the full suites pass 96 of 96 and 110 of 110 run one test at a
time. With the handler's window made unreachable and the ceiling lowered to 5
seconds, the test stops after 6 seconds, prints the count it reached, and
reports ERROR #8 with the harness recording a failure rather than a pass. With
the count raised past what the budget allows, a run that reaches two windows
still passes, so falling short and reaching nothing stay distinct. The
TX_NOT_INTERRUPTABLE branch, which no configuration in either suite builds, was
compile-checked in both copies with the configurations' own compile commands.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
c8d27c4e25 |
Stopped the SMP stack check analyzing a stack it just found broken (#648)
TX_THREAD_STACK_CHECK detects a broken stack, calls the error handler, and then
tests whether the word below the high-water mark still holds the fill pattern.
On the SMP side that second test is a plain if, so a thread whose stack has
just been reported as corrupt goes straight on into _tx_thread_stack_analyze().
Analyzing a stack that is known to be broken is what that function is least
able to do. It binary searches between stack_lowest and stack_highest for the
fill pattern and then scans forward with
while (*stack_ptr == TX_STACK_FILL)
which has no bound of its own and no reason to terminate once the pattern it is
looking for is no longer where the pointers say it should be. The non-SMP copy
was given an else for exactly this reason. The SMP copy never was, and the two
macros are otherwise identical, line for line, so this single keyword was the
whole of the divergence.
The path is live in CI rather than theoretical. Instrumenting the internal
handler and running all 110 binaries of stack_checking_build shows
threadx_thread_stack_checking_test reaching it four times per run, on a thread
whose stack the test corrupts on purpose. Every one of those four currently
falls through into the analyze it should be skipping.
This is not the timeout the SMP suite has been failing on.
threadx_thread_priority_change never reaches the error handler at all, so
whatever wedges it in teardown is something else. Worth closing regardless: a
runaway scan inside stack analysis would present as a test that stops producing
output and is eventually killed, which is the shape that has been costing this
suite whole runs, and it would be indistinguishable in the log from the hang
already being chased.
Verified on both configurations that define TX_ENABLE_STACK_CHECKING.
threadx_thread_stack_checking_test, the one test that exercises the changed
branch, passes 60 consecutive runs, and stack_checking_build and
stack_checking_rand_fill_build both pass 110 of 110, repeated at the
parallelism CI uses.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
fe40079353 |
Stopped the thread priority change test leaving core 0 to a finished thread (#647)
The SMP suite has been failing on threadx_thread_priority_change since 30 June.
The test reports SUCCESS and the process then never exits, so ctest kills it at
the thousand second timeout, twice, and the log carries nothing past the result
line. Instrumenting the harness teardown produced the state at the hang:
last stage reached: test_control_return: control thread resume returned
_tx_thread_preempt_disable: 0
core 0: current=thread 0 execute=thread 0
thread test control thread state=0 priority=0 threshold=0 core_control=1
thread thread 0 state=1 priority=0 threshold=0 inherit=0
Core 0 is held by a thread in state 1, TX_COMPLETED, while the control thread
sits in state 0, TX_READY, at priority 0. The Linux SMP scheduler fills a core
only when _tx_thread_current_ptr for it is null, and clears that pointer only
for a thread carrying a deferred preemption, which thread 0 is not. So core 0
can never be handed on, and with TX_THREAD_SMP_ONLY_CORE_0_DEFAULT and
TX_SMP_NOT_POSSIBLE the control thread has nowhere else to go. The scheduler
re-reads the same state every two milliseconds for as long as it is allowed to.
What put thread 0 at priority 0 is the last thing this test does:
thread_0.tx_thread_inherit_priority = 0;
_tx_thread_smp_simple_priority_change(&thread_0, 16);
with the stated intent of reaching the branch where the new priority is below
the inheritance priority. Zero cannot reach that branch, because 16 is not less
than 0. The other branch runs instead, and that branch assigns the inheritance
priority as the thread's priority while the code after it links the thread into
the list for the new priority regardless. Thread 0 therefore came away claiming
priority 0 while living in the priority 16 list.
Both halves of that hurt. Priority 0 ties with the control thread, so resuming
the control thread raised no preemption and left the execute pointer alone. The
mismatch between the recorded priority and the list the thread is linked into
then means that completing thread 0 removes it from a list it was never in,
leaving core 0 pointing at it for good.
Give the inheritance priority a value above the new one, which is what the
branch the comment names actually requires, and put it back to
TX_MAX_PRIORITIES afterwards so nothing downstream reasons about an
inheritance that is not there. Hold protection across the call as well: this is
an internal routine that expects it, and it was being called in the open.
Measured before and after by printing the thread's state at the point the test
reports success. With the inheritance priority at 0 it is priority 0 threshold
0, matching the hang above, on every run. With it above the new priority it is
priority 16 threshold 16, which agrees with the list the thread is linked into,
and resuming the control thread preempts core 0 the ordinary way.
All five SMP configurations pass 110 of 110 at the parallelism CI uses.
The comparison in _tx_thread_smp_simple_priority_change deserves a second look
on its own account. When its else branch runs, the thread's recorded priority
and the list it is linked into disagree by construction. Only this test is
known to reach that branch, by supplying an inheritance priority that cannot
arise in ordinary operation, so nothing here claims a defect in shipped paths.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
83dfbc3534 |
Made a teardown hang in the SMP suite say where it stopped (#646)
The SMP regression suite times out in CI on threadx_thread_priority_change and the log carries nothing that says why. The reason the log is empty is mechanical: test_control_return() opens with fflush(stdout), and that is the last flush before either the test finishes or it wedges. Everything printed after it sits in stdout's block buffer, and ctest discards that buffer when it kills the process at the timeout. So the log ends at the test's own result line no matter where the process actually stopped. That is enough to place the hang, if not to explain it. The failing runs print "SUCCESS!" and then stop, which means every check in the test body ran and passed, and the wedge is somewhere between that flush and exit(). The run of 30 June shows the same signature before any of this test's waits were bounded, so the hang is not the unbounded wait removed earlier, and the message added then for an exhausted cap never appears. Retrying tells us nothing new either: the suite spends 1000 seconds per attempt, twice, to reproduce the same silent timeout. Record how far teardown gets, and bound it. A stage variable is updated at each step from test_control_return() through test_control_cleanup() to exit(), and a watchdog thread armed on entry to test_control_return() reports the last stage reached, the per-core scheduler state, and every thread on the created list, then exits 99. The watchdog covers teardown and not the test body, because the test body has no bounded runtime to hold it to. Several tests here wait on a probabilistic interrupt window: threadx_thread_wait_abort_and_isr_test has been measured between 0.34 and 439 seconds while passing. Teardown is a fixed amount of work that takes milliseconds, so a bound on it cannot turn a slow pass into a failure. The default is 60 seconds, which also means a wedged run now reports in one minute rather than burning the 2000 seconds two 1000-second attempts cost today. The report is written with write() rather than printf() because a wedged thread may be holding the stdio lock, and a watchdog that blocked on that lock would reproduce the silent timeout it exists to replace. For the same reason it reads the ThreadX globals directly and takes no kernel lock; the values may be torn, which is acceptable for a post-mortem and cannot deadlock. One walk in test_control_cleanup() is bounded as well. The loop that steps past the timer thread and the control thread has no terminating condition of its own and spins for good if _tx_thread_created_count and the created list ever disagree, which is one of the shapes the timeout could be taking. It now reports and stops instead. Off by default in the sense that matters: stderr stays empty and stdout keeps its buffering, so output is unchanged on a passing run. TX_TEST_TEARDOWN_TIMEOUT overrides the bound in seconds and zero disables the watchdog; TX_TEST_TEARDOWN_TRACE echoes each stage as it is reached and line-buffers stdout so the surrounding output survives a kill too. Verified against an injected hang at the point the failing runs stop: the watchdog fires, names the stage, and exits 99. The dump is already informative, showing thread 0 left at priority 0 with threshold 0 and inherit 0, the same priority as the control thread, with core 0's execute pointer still on it. All five SMP configurations pass 110 of 110 at the parallelism CI uses, in both quiet and trace modes, and the suite runtime is unchanged. Only the SMP harness is instrumented. The non-SMP suite has not shown this hang. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a5483f0773 |
Bounded the wait for the delayed suspension window (#645)
threadx_thread_delayed_suspension_test waits for an interrupt to land while
thread 2 is part way through suspending, and waits for it with no bound:
while(delayed_suspend_set == 0)
{
tx_thread_wait_abort(&thread_2);
tx_thread_relinquish();
}
How long that takes depends on the build to a degree that is easy to miss. The
loop finishes in between a tenth of a second and three seconds in four of the
five ThreadX configurations. In trace_build it took 490 seconds, which was 40
percent of the whole ThreadX suite and more than every other test in that
configuration put together.
This is the third test in these suites built the same way, after
threadx_thread_priority_change and threadx_thread_wait_abort_and_isr_test: spin
until an interrupt happens to land in a narrow window, with nothing to stop the
spin if it does not. The other two have been given bounds already.
Give this one a wall clock budget too, for the same reason as the last: a tick
is delivered only when the port's timer thread runs, so the tick clock falls
behind real time under load or instrumentation, and instrumentation is exactly
what trace_build turns on.
The check after the loop needs care that the other two did not. It compares
thread_2_counter against thread_2_counter_capture, and the capture is taken
inside the interrupt handler at the moment the window is hit. Leaving that check
in place after a run that never reached the window would compare a live counter
against the zero it was initialised to and report a defect that is not there. So
the check is skipped when the window was not reached, and the run says so.
Reaching the window still exercises it exactly as before.
Verified both ways in trace_build, which is the configuration that was slow: the
window is reached in 8 seconds here and the test passes as it always did, and
with the budget forced to zero the test reports that the window was not reached
and passes without the dependent check firing.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
4d9ce41845 |
Gave the wait abort ISR test a budget instead of an open-ended wait (#644)
threadx_thread_wait_abort_and_isr_test waits for an interrupt to land while the
preempt disable flag is set, and waits for it ten times, twenty in the SMP copy,
with no bound on how long that takes. The handler in the same file already says
what can go wrong:
It is possible for this test to get into a resonance condition in which
the ISR never occurs while preemption is disabled
and perturbs its own duration to break out of it. That helps but guarantees
nothing, and if the resonance holds, the loop does not end.
It is also, by a wide margin, the most expensive thing in either suite. Run one
test at a time in CI it took between 148 and 726 seconds per configuration:
1936 seconds of the ThreadX suite's 2209, against 273 seconds for the other
ninety five tests together. Nothing else in the suite is within two orders of
magnitude of it.
The budget is in wall clock seconds, not ticks. That distinction turned out to
matter. A tick is delivered only when the port's timer thread gets to run, so
the simulated clock falls behind real time under load or under coverage
instrumentation, and never makes the loss up. A first attempt bounded the wait
at 20000 ticks, nominally 200 seconds, and failed to stop a run that took 726
seconds, because fewer than 20000 ticks had gone by. tx_time_get() cannot bound
elapsed time here; time() can.
A run that falls short says how many windows it reached rather than going quiet,
and the check after the loop is untouched. That check compares semaphore
bookkeeping which holds whatever number of windows were hit, so it still means
exactly what it did before. Hitting the race a few times rather than ten is a
smaller loss than it looks: the value is in reaching the window at all, and the
later hits repeat what the first ones established.
Verified by forcing the budget to 3 seconds, where the test stops after 3.14
seconds of wall clock and reports reaching 0 of 10 windows, with the following
check intact. At 120 seconds both suites pass every configuration run serially.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
7928dc4268 |
Ran the ThreadX and SMP regression suites one test at a time (#643)
A large part of both suites sleeps and then asserts on the tick counter, either
exactly or within a tick:
tx_thread_sleep(18);
now = tx_time_get();
if ((now == 18) || (now == 19))
The Linux port drives ticks from a thread running under SCHED_FIFO, so ticks
keep arriving whether or not the thread waiting on them can get a core. Starve
that thread and further ticks land between the sleep expiring and the read, the
value is past the one asked for, and a kernel that behaved correctly is reported
as broken. Thirty-five tests here and forty-one on the SMP side sleep and then
consult the clock or a counter driven by it, so this is most of the suite rather
than a corner of it.
The starvation is self-inflicted. Four tests are run at once on a four vCPU
runner, and each one is a process carrying a scheduler thread, a SCHED_FIFO
timer thread and a thread of its own, so the machine is oversubscribed two or
three times over by design. That is why three of these failed in the run of
18 August, and why the retry hid two of them.
Naming the sensitive tests and keeping them apart was tried first and does not
converge. A list covering the tests comparing tx_time_get() for equality missed
threadx_timer_multiple_accuracy_test, which compares timer-driven counters, and
CI failed on it. Widening the list to cover those missed
threadx_thread_sleep_for_100ticks_test, which asserts a range rather than an
equality, and CI failed on that. A list that is quietly incomplete is worse than
no list, because it reads as protection.
So stop overlapping the tests. Serial execution removes the contention for every
test at once, needs nothing to be enumerated, and makes a run reproducible: a
test either passes on its own machine or has a real defect.
The cost, measured in a four CPU cpuset, is close to a factor of four: the
ThreadX suite goes from 12.2 to 47.6 seconds for a configuration and the SMP
suite from 15.2 to 60.0 seconds. The suites parallelise almost perfectly, so
that factor is what parallelism was buying. It is worth giving up. The run this
replaces spent 36 minutes and reported a timeout carrying no information, and
2000 of those seconds went on retrying a test that had already hung twice.
Serial also makes the tick budget in threadx_thread_wait_abort_and_isr_test mean
what it says. Under contention that test took 255 seconds while its 20000 tick
budget never engaged, because the ticks themselves stretch when the process
cannot get a core. With nothing else running, ticks track wall clock and a
budget in ticks bounds elapsed time.
The FreeRTOS suite is left alone. It covers the creation paths of the
compatibility layer and has no tick accuracy tests, so it has nothing to gain
here.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
f2d27de25e |
Stopped a dead apt mirror from taking the whole install down with it (#642)
install.sh reaches the network four times, and on this runner pool that is not dependable. apt-get update stalled seven times in a single day: once for 55 minutes, once for more than two hours, and five times against the ten minute step timeout added alongside this. The log says the same thing every time. Every fetch from azure.archive.ubuntu.com comes back Ign, apt falls back to archive.ubuntu.com, and then the step produces no further output at all until something kills it. Nothing here bounded a fetch and nothing retried one, so a mirror being down cost a whole run instead of a few seconds. There is a second problem in the same lines. This script has no set -e, so a failed apt-get update did not stop the apt-get install that follows. The install went ahead against whatever package index the image happened to have, and the run failed later, somewhere with much less to say about why. Bound each attempt from outside and retry it. apt's own Acquire timeouts were tried first and are not enough: with them in place a run still sat inside a single apt-get update for nine and a half minutes without printing a line, having got as far as fetching noble-security InRelease. The retry loop never got a turn, because the first attempt never returned, and the step timeout was what eventually killed it. Whatever apt waits on there is not what Acquire::http::Timeout covers, so the bound has to come from outside the process. timeout does not care where the wait is. The Acquire options are kept anyway, since they make a slow mirror give up sooner, and pip gets its own retry and timeout flags for the same reason. timeout goes under sudo rather than over it, so that it signals apt itself. Signalling sudo risks the kill landing on sudo while apt carries on holding the dpkg lock, which would leave every retry failing for a different reason than the one being retried. The explicit exits stop a failed fetch being carried forward into a build. set -e is deliberately not used. rm -rf /opt/hostedtoolcache runs without sudo against a root owned directory and its exit status is not something this script should start depending on. The bounds fit inside the ten minute step timeout. Two minutes per attempt, three attempts, with 10 and 20 second backoffs, caps a command at about six and a half minutes, and a command that exhausts its attempts exits rather than letting the next one start. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
dafb7d70cc |
Stopped a stalled install from costing a whole regression run (#641)
The Install softwares step runs apt-get update, an apt-get install and a pip install, and it has stalled twice: 55 minutes on the SMP job of one run, and more than two hours on the ThreadX job of the next, against a normal 27 to 152 seconds across every other run measured. In the second case the tests never started at all. No step in this template had a timeout, so a stall runs until the six hour job limit. That turns a transient apt or PyPI problem into a lost run, and it hides what happened: the job simply sits there, and the failure that eventually gets reported says nothing about which step was stuck. Bound the three steps that do real work. Ten minutes for the install, against a normal worst case of 152 seconds. Fifteen for the build, which has run between 4 and 54 seconds. Sixty for the test step, which is the only one whose length depends on the suites themselves; the longest observed is 37 minutes, and that was with a wait in one test that has since been bounded. A step that trips its timeout fails and names itself, which is the point. None of these numbers is tight enough to trip on work that is merely slow. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b3486f9b46 |
Bounded the wait in the thread priority change regression tests (#640)
Both copies of this test, the SMP one and the non-SMP one, install an interrupt
handler and then spin until it clears a flag:
test_isr_dispatch = test_isr;
do
{
...
} while (test_isr_dispatch);
The handler clears that flag only on a narrow window: thread 3 at priority 6,
ready, and not yet at the head of its priority list, which exists only part way
through a priority change. If an interrupt never lands inside that window, the
loop never ends.
That is what has been failing in CI. The SMP suite has been red since 30 June,
and this test times out in three of the last four failing runs, always in a
stack-checking configuration. The evidence that it is a hang rather than slow
work: the test carries no per-test timeout property, so ctest's --timeout 1000
applies, and locally the test finishes in 0.12 seconds with a worst case of 0.29
over thirty runs. Nothing turns that into more than a thousand seconds. It also
survived --repeat until-pass:2, so it hung twice in succession.
The TX_NOT_INTERRUPTABLE path in the same handler already stops after a fixed
amount of work. Only the interruptable path, which is the one the failing
configuration uses, had no protection.
Cap the loop and clear the handler on the way out. When the window is not reached
the test says so and still passes: not reaching it is a gap in what this run
covered, not a fault in the code under test, and failing would report a defect
that does not exist. The counters are left alone so the checks that follow keep
their previous meaning.
The cap is 100000 attempts. An exhausted cap takes about 60 seconds, measured,
and a successful run takes 0.12 seconds, which puts the usual cost around two
hundred attempts and leaves the cap roughly two orders of magnitude clear of it.
That is wide enough not to lose coverage on a slower machine, while replacing a
timeout that says nothing with a message that says what happened.
Not reproducible here, which fits the diagnosis rather than contradicting it: on
sixteen cores the window is hit almost at once. Thirty sequential runs, two
hundred at parallelism thirty-two, sixty pinned to two CPUs, forty pinned to one,
and three full-suite passes at the parallelism CI uses all came back clean. The
defect is the reliance on the window, not any particular machine.
Verified with the window deliberately made unreachable: before this change the
test runs until it is killed, and after it exits in about a minute reporting that
the window was not reached. The full SMP suite passes 110 of 110 at CI's
parallelism.
|
||
|
|
ca62edd27d |
Swept the stack-heavy measurement across placements, and qualified its result (#638)
#636 reported that a stack in BTCM gave a threefold tighter spread than DRAM0 for stack-heavy work. That measurement used a single code placement, which is the methodology #631 and #633 exist to correct: the cache benchmark got an alignment sweep and the interrupt handler got one, and this measurement never did. It was noticed when #637 added two threads to the same image and the figure moved -- both spreads came out near 6500 and the minima rose 15%. The recursive body is now generated at four placements and all four are measured, per placement, in one image. placement BTCM min / spread DRAM0 min / spread offset 0 41854 / 6850 42036 / 6880 offset 16 48670 / 1946 48792 / 1978 offset 32 42388 / 6914 42752 / 7018 offset 48 48914 / 1860 49198 / 6786 Reproducible across runs to within a few hundred cycles. Spread is dominated by code placement rather than by the memory holding the stack. It ranges from 1860 to 6914 depending on where the body falls in a cache line, and placement also moves the minimum by 17%, from 41854 to 49214. Against that, the memory contributes a consistent but small advantage: BTCM's minimum is lower at all four placements, by 0.4% to 0.9%. BTCM's spread beats DRAM0's decisively at one placement of the four, offset 48, at 1860 against 6786. At the other three the two are within 2% of each other. So the effect #636 reported is real where it occurs and is not a property of the part: quoting it as one invited the reader to expect it everywhere. #636's claim should be read as qualified by this. A stack in BTCM buys a small consistent improvement in the best case and a large improvement in spread at some code placements and not others. Anyone building a determinism argument on it needs the placement sweep in the loop, not a single figure. The pad nops that displace each placement execute on every recursion level rather than once, so each placement carries a slightly different constant cost, about 0.6% at the widest. That cancels in the BTCM against DRAM0 comparison, which is made at the same placement, and does not affect spread within one. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
1ef49338e9 |
Gave each thread its own MPU window, and made a violation fault (#637)
First step towards a ThreadX module port for this core: establish that PMSAv8-R
regions can be switched per thread on this part, what that costs, and that a
violation actually faults. Those are the questions worth answering before
writing a module manager on top of them.
Two threads each own a 4 KB window at the top of DRAM2. The windows are carved
out of the broad data region in mpu.c, because isolation is only meaningful in
memory no other region already covers -- every other region in that map is a
wide RW window, so a private buffer inside one of them would be reachable by
every thread whatever else was programmed.
Each thread writes its own window, which must succeed, and then reaches for the
other thread's, which must fault. The second half is the part that matters: a
test that only shows a thread reaching its own memory would pass just as well
with no protection at all.
Measured on the S32Z280-594EVB, reproducible across three runs:
thread 0 window 0x3187E000 own: reachable other: faulted
thread 1 window 0x3187F000 own: reachable other: faulted
region switch cost: 562 to 604 cycles
The cost is worth noting for the module port to come. A context switch on this
part is about 1400 cycles, so switching one region adds roughly 40% to it, and
most of that is the dsb and isb rather than the register writes. A module switch
programming several regions should therefore batch the barriers once at the end
rather than per region.
Scope, stated plainly. The window is applied by the thread calling
thread_mpu_activate, not by the scheduler. The port's scheduler does call
_tx_execution_thread_enter under TX_ENABLE_EXECUTION_CHANGE_NOTIFY, which would
make it automatic, but that macro is read by port assembly compiled into the
shared threadx library, so enabling it would oblige all nine example targets in
this port to supply the four execution hooks. A ThreadX module port carries its
own copies of the port assembly for exactly that reason, and that is where the
switch belongs. There is no user mode, no syscall boundary and no loader here.
The fault is survivable the same way the boot probes make it survivable:
fault_expected tells the data abort handler to record the violation and resume
after the faulting access. That works in thread context because the handler
returns where it came from rather than to a fixed recovery point.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
f535a4ec67 |
Measured stack-heavy work against the memory holding the stack (#636)
#635 found BTCM worth about 7.4% on a context switch with no determinism advantage, and said why: a switch saves sixteen registers, roughly one cache line, so the stack's cache state has almost nothing to contribute. It named the interesting case as work with a large stack working set, said it had not been measured, and said it should not be assumed. This measures it, and the answer inverts the earlier one. deep_touch recurses 24 frames, writing a frame on the way down and reading it on the way up, so the working set is the whole descent. The cache is cleaned and invalidated before each sample, so every descent starts cold. Two threads, one stack in BTCM and one in DRAM0, no partner threads and no relinquish: the timed region is entirely within one thread. Reproducible across runs: min mean max spread stack in BTCM 42868 43031 44890 2020 stack in DRAM0 42982 43291 49842 6860 The mean is the same to within 0.6%. The worst case is 10% lower for BTCM and the spread is 3.4 times tighter. No sample in either configuration exceeded twice the minimum, so these maxima are the workload rather than a timer tick -- which is the mistake that produced a false jitter result in #635 and is why the count of interrupted samples is printed. Put beside #635 the two measurements say opposite things and both are true. For a context switch, a small footprint touched every time, BTCM buys throughput and no determinism. For stack-heavy work, a large footprint touched once, it buys determinism and almost no throughput. The reason is that this workload is compute bound at the optimisation level this BSP builds at: 43000 cycles for 24 frames is dominated by call and loop overhead, so line fills are a few percent of the total and barely move the mean. What they do is vary, and that variance is what a bank with no cache in the path removes. So the determinism argument for TCM holds here, but it is worth 10% of worst case and a threefold narrowing of spread, not an order of magnitude. Anyone citing this in a safety argument should cite those numbers and not a larger claim. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
367d91880b |
Measured context-switch cost against the memory holding the stack (#635)
#634 placed a thread stack in BTCM and deliberately claimed no timing benefit, because none had been measured. This measures it. Two pairs of equal-priority threads hand control back and forth with tx_thread_relinquish. One pair has both stacks in BTCM, the other in DRAM0, and the measuring thread of each pair times the round trip in PMU cycles. Both pairs run in one image from one copy of the measuring code, which is what makes the comparison safe: the alignment trap that invalidated earlier work here bites when two builds with different layouts are compared, and a code shift moves both pairs equally. Reproducible to the cycle across runs: min mean max BTCM stacks 1370 1379 1402 DRAM0 stacks 1476 1480 1508 BTCM is about 7.4% faster, or 110 cycles on a round trip of two switches. Three findings that bound the claim, and the last one deflates it. The figure holds whether the cache is warm or cold. Cleaning and invalidating the data cache before every timed switch costs both configurations about 40 cycles and leaves the gap at 7.4%: warm it is 1334 against 1440, cold 1370 against 1476. So the advantage comes from BTCM's zero wait states, not from avoiding cache misses. That is because a context switch touches almost no stack -- sixteen registers, about one cache line -- so the stack's cache state has little to contribute either way. TCM should matter much more for threads with deep call chains or large locals, where the stack working set is big enough for cache state to dominate. That is not measured here and should not be assumed. There is no determinism benefit visible in this test. Excluding preempted samples, jitter is 32 cycles for BTCM and 28 to 34 for DRAM0 -- comparable, not better. A first version of this measurement appeared to show BTCM with 16 times less jitter, and that was wrong: max was reporting whichever pair a timer tick had landed on. Across three runs the outlier appeared in the BTCM pair once and the DRAM0 pair twice. Samples past 2000 cycles are now counted separately and excluded from min, mean and max alike, and the count is printed so the reader can see how many there were. Also tried and discarded: loading the partner thread with a cache walk to create pressure. The timed round trip includes the partner, so the walk dominated every sample and put all 256 past the outlier threshold. The per-sample flush replaced it and sits outside the timestamps. The demo also starts the PMU cycle counter, which bsp_boot.c does for the probe image and this image never ran. Without it every reading would have been zero, which reads as a free context switch rather than as a dead counter. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
62e966f6bf |
Enabled BTCM, measured it, and put a ThreadX thread stack in it (#634)
* Enabled BTCM and measured what it offers as a data store
BTCM was disabled because of a regression that turned out not to exist; the
claim was retracted in the previous commit. It is enabled now, and this
measures why that is worth doing: 16 KB at zero wait states, where ATCM has
one, and no cache in the path at all.
Enabling it needs three things, all of which existed for ATCM already: the
region register write at EL2, an MPU region, and an ECC preload before any read
(TRM 6.2.2). BTCM accepts 32-bit stores where ATCM needs 64-bit, which
tcm_preload already handles.
The cache benchmark now sweeps three memories rather than one, four loop
alignments each. At the alignments where the loop is not instruction-fetch
bound:
memory cold (uncached) warm (cached) gain
DRAM2 half-speed 857,540 651,436 24.0%
DRAM0 full-speed 797,824 651,297 18.3%
BTCM zero wait 694,689 651,369 6.2%
Three things follow.
Warm times are identical across all three memories, within 0.02%. Once the data
cache is working the backing store barely matters, because the working set fits
in it.
Cold times rank as the reference manual predicts: BTCM fastest, then DRAM0,
then DRAM2 at half the core frequency (S32Z2 RM 6.3.6).
BTCM still shows a 6.2% gain when the caches are enabled, and that cannot be
the data cache, because an enabled TCM is Non-cacheable Non-shareable Normal
memory whatever the MPU says. It is the instruction cache on the timing loop.
This probe has always measured both caches together; three memories side by
side is what makes that visible.
The number that matters for placing data in BTCM: uncached BTCM is within 6.6%
of the best cached case, where uncached DRAM0 is 22% off it. Data in BTCM runs
at close to cache-hit speed with no cache to miss, which is the determinism
argument stated as a measurement rather than an assertion.
DRAM2's sweep is unchanged with BTCM enabled, 0 and 0 and 240 and 240, which
independently confirms the retraction.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Put a ThreadX thread stack in BTCM
The code side of TCM was done in #630; this is the data side. One of the demo's
three thread stacks now lives in BTCM and the other two stay in DRAM0, so a run
exercises both paths and a mistake in either shows up.
link.lds gains a BTCM region and a .btcm_bss NOLOAD section, so a stack is an
ordinary C array with a section attribute and the linker checks it fits, rather
than a hardcoded address that silently overflows the bank.
entry.S preloads the whole bank at EL2, and that is not optional. ECC is enabled
on this part, so a TCM location must be written before it can be read (TRM
6.2.2), and a stack is read before the program writes it -- the first context
restore pops what tx_thread_create built into it. The preload has to happen
before any C runs, because the demo images do not run bsp_boot.c, which is where
the ATCM preload lives. 32-bit stores suffice for BTCM where ATCM needs 64-bit.
Why BTCM for a stack: 16 KB at zero wait states where ATCM has one, and never
cached whatever the MPU says about it. Measured in the previous commit, uncached
BTCM comes within 6.6% of the best cached case while uncached DRAM0 is 22% off
it, so stack access runs at close to cache-hit speed without depending on a line
being resident. That is the property a determinism argument needs.
What this commit does not claim: no thread-level timing improvement has been
measured. The case for BTCM here rests on the memory characterisation and on
removing the cache from the path, not on a measured context-switch figure. That
measurement is worth doing and has not been done.
Verified on the S32Z280-594EVB. The demo reports its stack addresses so the
placement is visible rather than implied -- sleeper at 0x30100000 in BTCM,
spinner and judge in DRAM0 -- and passes with 100 ticks, 20 sleeper wakeups and
20 preemptions, so a real thread schedules, preempts and context-switches on a
tightly-coupled-memory stack. The boot image still passes six of six probes.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
35de56853c |
Retracted the claim that a second TCM bank costs the data cache (#633)
entry.S and readme_s32z280.txt both stated that enabling any second TCM bank removes all measurable data-cache benefit on this part, and gave five configurations as evidence: ATCM alone at a 24% cache gain, four combinations involving a second bank at none. Both concluded the cause was not a particular bank, not its address and not the ECC preload, but enabling a second bank at all. Both cited the Cortex-R52 and S32Z2 errata as not covering it and offered a shared RAM pool between the LLC and the TCMs as an explanation. A defect report went to NXP on that basis. It was an artifact of the benchmark. That benchmark was bimodal with respect to where its timing loop fell inside a 64-byte cache line, reporting either 24% or nothing at all for identical silicon, and every one of those five configurations was an edit to entry.S, so every one shifted the code that followed and moved the loop between modes. Adding two nop instructions reproduces the "second bank" figure exactly, to the digit. The report to NXP has been withdrawn. Re-measured with the alignment sweep added in #631, one bank and two are indistinguishable: loop offset in line ATCM only ATCM + CTCM 0 gain 0 gain 0 16 gain 0 gain 0 32 gain 240/1000 gain 240/1000 48 gain 240/1000 gain 240/1000 So enabling a second bank costs nothing measurable. The banks stay disabled, but for the ordinary reason that nothing in this example uses them, and both texts now say that instead. Enabling one is a single line, with the ECC preload before any read (TRM 6.2.2) and an MPU region as the only prerequisites, both already handled for ATCM. The readme also now states the general point, which outlasts the TCM detail: a single-figure timing result from this example cannot be compared across builds unless the timed loop's alignment is controlled, because almost any change shifts code. Comments and documentation only; no generated code changes. Verified on the board regardless, since entry.S was touched: six of six probes pass and the sweep is unchanged. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
3aaaa6700d |
Revalidated the ATCM handler result across four placements, and it holds (#632)
The handler comparison in #630 measured one alignment, which is the mistake that made the cache benchmark in this example report 24% or 0% for identical silicon. The handler body is now generated at four offsets within a cache line, all four are measured, and the figures are reported per placement. The claim survives. Mean cycles for the handler body: loop offset code RAM ATCM 0 523 383 16 534 377 32 539 375 48 521 383 ATCM is faster at every placement, by about 28%, and the two sets of means do not overlap. Worst case improves as well, 510 against 694. Two things worth recording beyond the headline. The handler measurement is only mildly alignment sensitive, 3.5% across placements in code RAM and 2% in ATCM, quite unlike the cache loop's two modes. So this comparison was less fragile than the cache one, and #630's direction was right even though its method was not defensible. The absolute numbers differ from #630 because the body now sits behind a placement wrapper that adds a call; the comparison is internally consistent either way. Both variants also report identical cache sweeps, 0 and 0 and 240 and 240, which settles the regression this branch's predecessor appeared to show. That apparent regression was the single-alignment probe moving between its two modes, not anything about ATCM. One copy of the logic is kept: the wrappers inline a single always_inline implementation, so the four placements cannot drift apart. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
1d3f3f8e4c |
Measured the cache benchmark at four alignments, because one is not enough (#631)
This benchmark was bimodal and reported a single number, which made it worse
than no benchmark. The same workload on the same silicon reports either 24%
cache benefit or none at all, decided only by where the loop falls inside a
64-byte line -- and therefore by any unrelated change that shifts code
ahead of it. Two nop instructions added to entry.S were enough to flip it.
That is not a hypothetical. A run of conclusions drawn from this probe turned
out to be measuring code layout: an interrupt-handler comparison, a claim that
enabling a second TCM bank costs all cache benefit, and a follow-up claim that
what mattered was when the TCM region register was written rather than what it
contained. The last of those was reported to NXP as a defect and has had to be
withdrawn. Enabling CTCM and adding two nops produce identical results, to the
digit, because the only measurable consequence of the enable was the eight
bytes of instructions it added.
Four copies of the loop are now generated at different offsets within a cache
line, all four are measured, and the low and high gains are both reported.
Pinning a single alignment was tried first and is not a fix: it silently picks
one of the two modes -- aligned to 64 the loop sits permanently in the low one.
Measured on the S32Z280-594EVB, reproducing exactly across runs:
loop offset in line cold warm gain
0 890,302 890,035 0%
16 890,208 889,976 0%
32 857,439 651,390 24.0%
48 857,631 651,571 24.0%
The cold pass differs between the modes as well, 890k against 857k, so the
loop is slower even with both caches off. The cold pass is instruction-fetch
bound out of code RAM at half the core frequency (S32Z2 RM 6.3.6), and how the
loop straddles lines decides how much of the data cache's contribution is
visible at all. This probe therefore measures both caches together and always
did; the sweep at least makes the variation visible instead of letting one
arbitrary placement stand in for the part.
C4 now passes if any alignment shows a 10% speedup, and says so explicitly
when the low mode does not, so the sensitivity appears in the log rather than
being discovered later.
Verified: with the sweep in place, adding 0, 8, 12 or 20 bytes of nops to
entry.S leaves the reported low and high gains unchanged. Before it, the same
shifts read 24.0%, 0%, 0% and 0%.
Also adds cache_disable_all, which the sweep needs: cache_enable was one-way,
so a second cold reading in one run was impossible.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
b2983c08d3 |
Ran the interrupt handler from ATCM, and measured what that buys (#630)
The TCM work so far enabled ATCM and left it empty, which buys nothing.
This places code in it and measures the result.
link.lds gains an ATCM region and an .atcm_text section whose run address
is in the bank and whose load address is in CODE. tcm_copy_atcm_text
moves it, using 64-bit stores because ECC is enabled on this part and
ATCM requires them (Cortex-R52 TRM 6.2.2); both ends of the section are
8-byte aligned so there is no narrower tail to leave without check bits.
The copy runs after T4 and T5, which write test patterns to the first and
last words of the bank and would otherwise land on top of the code.
s32z280_atcm.elf is the same image as s32z280_boot.elf with the interrupt
service body placed in ATCM. Both targets exist so the comparison can be
repeated on one board in one session without reconfiguring. The service
routine is split into a timed wrapper that stays in .text and a body that
moves, so the wrapper's own cost appears in both measurements and cancels.
Measured in PMU cycles over 64 samples, caches enabled in both:
code RAM ATCM change
min 454 334 -26.4%
mean 458 340 -25.8%
max 612 466 -23.9%
spread 158 132 -16.5%
CNTPCT is not used for this: at 8 MHz it cannot resolve a handler body,
let alone the variation in one.
The level shift is the solid part. ATCM is a quarter faster even though
the caches were on and code RAM had the instruction cache available,
which says the handler does not stay resident between interrupts 10 ms
apart -- so each one pays a cold fetch from code RAM, which runs at half
the core frequency where ATCM runs at full speed with one wait state
(S32Z2 RM 6.3.6).
The determinism claim deserves less weight than the numbers first
suggest. The spread narrows by only 16%, and ATCM's worst case still sits
slightly above code RAM's best case, so the two distributions overlap at
the tails rather than separating. Whatever jitter remains is not
dominated by instruction fetch.
Both images pass six of six boot probes.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|