* Compiled the module manager C sources, which no check had ever built
#672 corrected this script's assembly glob and brought the module ports into
the count, but only their assembly. Their C stayed outside every check: 27
files of portable module manager under common_modules, plus the per-port code
under ports_module/<core>/gnu/module_manager/src. By this script's own standard
-- a port simply absent from the count reads as covered -- 293 files across
nine Arm module ports were compiled by nothing, with either compiler.
Each module port ships its own tx_port.h and txm_module_port.h carrying the
control-block extensions the dispatch code needs, so a port is compiled against
its own headers rather than the base port's. Two details the ports themselves
dictate:
An SMP port's control blocks come from common_smp. Pairing cortex_a35_smp
with the single-core headers hid _tx_thread_smp_protect and
_tx_thread_smp_unprotect behind implicit declarations and lost
tx_thread_smp_core_executing from TX_THREAD, so fourteen files reported
errors for a port that builds correctly.
The TrustZone ports carry cmse_nonsecure_entry, which needs -mcmse to be
honoured rather than ignored. tx_thread_secure_stack.c also carries GCC's
optimize attribute, which clang does not implement; that divergence is
suppressed by name, for a file GCC builds cleanly.
All nine ports compile: 293 of 293. Verified that the stage fails as intended
by injecting a defect into a throwaway worktree -- a defect in common_modules
is reported under every port, one in a port's own source only under that port.
Assisted-by: Claude Code (Opus 5)
* Ran the module manager stage against the sources that reach it
Three corrections to the new stage, all found by running it against dev
rather than against the tree it was written on.
The workflows did not trigger on common_modules. Both check_clang.yml and
check_gcc.yml list ports_module but not common_modules, and the portable
module manager under it is the larger half of what the stage compiles --
28 of the 31 to 37 files each port builds, plus two include directories.
A change there left the stage unrun, which is the same "absent from the
count reads as covered" that the stage exists to close. Both lists gain
common_modules, and they stay identical to each other as the comment in
each asks.
A deliberate deprecation notice read as a build failure. Since this PR
was opened, txm_module_manager_absolute_load.c gained a #pragma message
steering callers to the extended entry point. The stage treats any
compiler output as a failure, so that one notice failed every port: nine
failures on a tree where nothing is wrong. Pragma messages are now
waived for the stage, because a notice to callers is not a defect in the
file that carries it.
That failure also printed nothing. Both C stages report by grepping the
output for "error:", so a diagnostic that is not an error produced a bare
FAIL line with no reason under it, and the only way to learn the reason
was to reproduce the compile by hand. Both stages now fall back to
showing what the compiler actually said.
The TrustZone attribute waiver is narrowed to the one file that needs it.
tx_thread_secure_stack.c carries GCC's optimize attribute, which clang
does not implement; it is the only file among the 300-odd this stage
compiles that does. Waiving the warning for the whole port would have
swallowed a stray unknown attribute anywhere else in it.
Verified with the same toolchain CI uses, ATfE 22.1.0: the full script
passes, and every port compiles every file.
cortex_a35 31/31 cortex_a35_smp 31/31 cortex_a7 34/34
cortex_m0+ 33/33 cortex_m23 37/37 cortex_m3 33/33
cortex_m33 37/37 cortex_m4 33/33 cortex_m7 33/33
The counts are each one higher than this PR first reported, because
txm_module_manager_absolute_load_extended.c has landed since. The stage
picked it up with no edit, which is what globbing the directories was for.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
---------
Co-authored-by: r <r@r>
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
Against the merged SMP coverage report -- every build configuration instrumented
and unioned -- sixty-four lines of common_smp/src were uncovered, 5114 of 5178.
Fifty-three of them are closed here and the report reads 5167 of 5178. The SMP
coverage floor goes from 98 to 99 with it.
Thirty-six of the sixty-four were one loop repeated four times: the walk in
tx_block_pool_delete, tx_byte_pool_delete, tx_event_flags_delete and
tx_queue_delete that releases every thread suspended on the object with
TX_DELETED. The suite deletes all four object types after every single test, and
that is exactly why the loop never ran. test_control_cleanup in the ThreadX
suite deletes the application's objects first and its threads last, so a test
that ends with a thread parked on a queue has that thread walked out of it by
tx_queue_delete. The SMP suite's cleanup deletes the threads first, deliberately
-- it was changed so that no application-owned object is still referenced when
the object loops run, which is what stopped a class of teardown hang. The side
effect is that all four deletes now run against an empty suspension list.
tx_semaphore_delete is the one member of the family that was already covered,
because threadx_semaphore_delete_test deletes a busy semaphore on purpose.
threadx_object_delete_suspension_test is the same idea for the other four. Two
threads suspend on each of a block pool, a byte pool, an event flags group and a
queue; the control thread waits on the object's own suspended count through
tx_*_info_get rather than on an ordering it cannot guarantee across four cores,
deletes the object, and checks both waiters came out with TX_DELETED. Two
waiters rather than one so the loop takes its back edge as well as its body, and
every wait is bounded in ticks so a suspension that never arrives fails the test
instead of hanging it.
threadx_trace_entry_update_test and threadx_thread_misaligned_stack_test are
ports of the two tests that closed the equivalent gaps in common/src, and close
fourteen more lines here: tx_block_allocate 123, 175, 182, 319 and 326,
tx_byte_allocate 130, 210, 217, 359 and 366, tx_thread_system_suspend 504 and
560, tx_trace_object_register 221, and tx_thread_create 133. The one substantive
change is core confinement. The trace test needs thread 0 to suspend and thread
1 to then release what it waits for; on four cores thread 1 gives the block back
before thread 0 has suspended and the update block behind the suspension is
never reached, so both threads are excluded from cores 1 to 3. The misaligned
stack test needed no such change.
threadx_byte_memory_long_search_test closes three of the eleven in
tx_byte_pool_search. Lines 264, 267 and 270 are the
TX_BYTE_POOL_MULTIPLE_BLOCK_SEARCH limit -- twenty on this port -- where a long
search drops and retakes protection so that it cannot lock the other cores out
for the whole walk. No byte pool in the suite ever had twenty fragments. This
one is filled with small chunks until it refuses another and then has every
second chunk released, so the free fragments are never adjacent and cannot be
merged, and the request is larger than any of them but smaller than the pool's
theoretical total, which is what makes _tx_byte_pool_search walk rather than
refuse at the door. The layout is asserted rather than assumed: the test checks
the fragment count and checks the probe request really does fail before the
workers start, because either would otherwise turn it into a silent no-op.
Eleven lines remain and they are not a to-do list. Eight are the delay loop in
tx_byte_pool_search that fires when another thread claims the pool inside the
window the search opens. The Linux SMP port serialises all four cores on one
pthread mutex, so that window is an unlock immediately followed by a lock on
that mutex, and glibc hands an uncontended mutex straight back to the thread
that just released it: measured over 180,003 windows across three cores, with
zero handovers. The shipped test therefore does 250 searches per worker rather
than the sixty thousand that probe used, because the twenty-block threshold is
crossed by the first search. The other three are in tx_thread_smp_utilities.
Line 149 is a range guard placed after the shift it is meant to guard, so
reaching it needs a shift by the width of the type; the fix is to move the check
above the shift, matching the TX_MAX_PRIORITIES > 32 variant of the same
function, and that belongs in its own change. Lines 1073 and 1074 need a mutex
owner that is genuinely executing on another core when a waiter suspends, and
three shapes were tried without producing one on this port.
Measured twice before and twice after, every gcda deleted between runs and
570 of 570 tests passing each time: 5114 of 5178 both times before, 5167 of 5178
both times after. Branch coverage goes from 2768 to 2821 and 2823 of 3548. A
floor of 99 needs 5127, so the ratchet lands with forty lines of headroom
against a numerator that has been seen moving by two between runs.
Assisted-by: Claude Opus 5 <noreply@anthropic.com>
* Assembled the module ports, which no check had ever compiled
scripts/check_clang.sh globbed ports_module/*/gnu/src, which does not exist --
the module ports keep their assembly in module_manager/src. The [ -d ] guard
skipped it in silence, so 116 assembly files across nine Arm module ports were
assembled by no check, with either compiler, in the script whose own comments
state three times that "a port that is simply absent from the count reads as
covered". Stage 1 goes from 724 of 724 to 840 of 840; the feature-macro stage
had the same gap and goes from 412 files to 469.
Correcting the path exposed five defects, and only one of them was a build
failure. The other four assembled cleanly and did the wrong thing, because GAS
runs the C preprocessor on .S and not on .s:
ports_smp/cortex_a7_smp/gnu/src/tx_thread_smp_unprotect.s, the only .s in a
directory of twenty-one .S, ignored all four of its own feature macros. It
wrote the caller's LR into the protection structure on every unprotect -- a
store guarded by TX_MPCORE_DEBUG_ENABLE -- sent an unconditional SEV, and
returned through both BX lr and MOV pc, lr. Its cortex_a5_smp and
cortex_a9_smp siblings are .S.
ports_module/cortex_m33/.../tx_thread_stack_build.s emitted both arms of an
#ifdef TX_SINGLE_MODE_SECURE, so the non-secure LR value overwrote the secure
one and the secure build got the wrong frame.
ports_module/cortex_m23/.../tx_thread_context_{save,restore}.S carried the
POP {r0, lr} that check_clang.sh's own comment describes as the reason the
feature-macro stage exists. The 16-bit Thumb POP takes r0-r7 and pc only.
The identical fix already sits in ports/cortex_m23/gnu/src; the module copy
never got it because nothing scanned it.
ports_module/cortex_m23/.../tx_thread_secure_stack_initialize.S used MOV
rather than MOVS for an 8-bit immediate, latent behind TX_SINGLE_MODE_SECURE.
Both siblings in the same directory already use MOVS.
ports_module/cortex_a7/gnu/module_manager/src is the one that failed to
assemble, on GCC 14.3 as well as on LLVM: #define SYS_MODE was never
expanded, so #SYS_MODE reached the assembler as an undefined symbol.
Twenty-nine .s files under gnu trees are renamed to .S. Every one of them is
already named .S by the build scripts that compile it, so this repairs those
scripts rather than churning them -- ports_module/cortex_a7's build_threadx.bat
names all eighteen with a capital S, and works today only on a case-insensitive
filesystem. Renaming rather than converting the #defines to GNU assignments is
what fixes the #ifdef blocks as well as the constants; the assignments would
have fixed two files and left twenty-seven silently ignoring their macros.
Files with no preprocessor directives are left as .s: they are not broken, and
check_ports.sh gains a check that keeps them that way. Only the gnu trees are
checked there -- the IAR, Arm Compiler 5 and Keil assemblers preprocess .s
themselves, and about three hundred files in this repository rely on that.
Verified with both toolchains on the same tree: 840 of 840 assembled by
ATfE 22.1.0 and by arm-gnu-toolchain 14.3.rel1, all five stages of
check_clang.sh green, and check_ports.sh green including the reproducibility
check. The new check was shown to fail by planting a copy of the file it was
written for.
No regression test accompanies this. The assembly it covers is executed by no
host test, and the check itself going from 724 files to 840 is the coverage
AGENTS.md asks for -- together with the new check_ports.sh section, which is
what stops the class recurring.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Fixed the AArch64 samples, none of which had ever linked with GCC
Every AArch64 gnu example build failed at the sample link, all 27 of them --
13 under ports/ and 14 under ports_smp/:
libg.a(libc_a-init.o): in function `__libc_init_array':
undefined reference to `_init'
relocation truncated to fit: R_AARCH64_CALL26 against undefined
symbol `_init'
libg.a(libc_a-fini.o): in function `__libc_fini_array':
undefined reference to `_fini'
build_threadx_sample.sh links with -nostartfiles, which is correct for a port
carrying its own reset path, and that drops crti.o and crtn.o along with
everything else. startup.S calls __libc_init_array by design, and newlib's
implementation calls _init, which crti.o is what defines. The AArch32 scripts
are unaffected: they use nosys.specs and never reach __libc_init_array.
The fix links crti.o and crtn.o explicitly, bracketing the object list -- the
first must precede every .init contribution and the second must follow all of
them, so their position is load-bearing rather than stylistic. Both paths come
from the compiler's own -print-file-name, so nothing here hard-codes a
toolchain layout.
The atfe branch sets both to empty, deliberately: picolibc's __libc_init_array
does not call _init, those 27 images link today, and adding crti.o would change
a working link for no reason. That is also why check_clang.sh is green on these
and does not list them as expected to fail -- the LLVM path never reached the
gap, so nothing has ever linked them and failed.
Fixed in ports_arch/ARMv8-A/threadx/ports/gnu/example_build, which is the
single source for both the ports/ and ports_smp/ copies, then regenerated with
update.sh --port-sets tx,tx_smp. The 27 generated copies are in this commit
because ports_arch_check compares them.
Verified: all 27 link with arm-gnu-toolchain 14.3.rel1 aarch64-none-elf, where
0 of 27 did before; _init and _fini disassemble to the expected crti prologue
and crtn epilogue over a ret; check_clang.sh with ATfE 22.1.0 is still green on
all five stages, including the 42 script-driven example builds; check_ports.sh
is green including the reproducibility check.
No regression test: these are link-only example images that no host test
executes. What guards them is check_clang.sh's example stage today, and
check_gcc.sh's, which is the next change and is the reason this was found.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Added a GCC check for the Arm ports, which nothing had ever compiled
GCC is the project's declared default compiler (AGENTS.md, "The default
compiler for the project is GCC 14 on Linux"), it is what the gnu ports exist
for, and it is what nearly every downstream user builds with -- and nothing in
CI compiled a line of any port with it. The only cross-compilation check that
ran was the LLVM one, so the ATfE path was better guarded than the GNU one, on
ports whose directory is literally named gnu. ci_cortex_m covers four port
families; this covers forty.
Five stages, mirroring scripts/check_clang.sh stage for stage:
1. assemble every .S and .s of every Arm gnu port -- 840 files
2. assemble again the parts behind TX_ENABLE_VFP_SUPPORT,
TX_ENABLE_FIQ_SUPPORT, TX_LOW_POWER and
TX_ENABLE_EXECUTION_CHANGE_NOTIFY -- 469 files
3. compile common/src for one core per architecture profile -- 185 x 9
4. link the script-driven example builds -- 42
5. link the CMake-driven Cortex-R52 images -- 5
Two scripts rather than one with a --toolchain flag: the flag surface differs
(a prefixed driver against --target=), the C library differs, and the set of
examples that can link differs. Folding them together makes it easy to weaken
one check while working on the other.
Two toolchains, both required. Arm ships arm-none-eabi and aarch64-none-elf as
separate downloads and PORT_TARGET maps every port to one of exactly those two
triples, so --arm-none-eabi and --aarch64-none-elf each take a driver or the
directory holding it, defaulting to the environment and then to PATH. A missing
one is a hard error rather than a soft skip: letting a run cover half the tree
and still report "all checks passed" is the failure this script exists to end.
PORT_TARGET is copied verbatim from check_clang.sh, including its warning not
to prefix-match core names -- cortex_a5* also matches the AArch64 cortex_a53.
VFP_EXTRA is the one map that is not a copy, and check_clang.sh's comment about
it is false for GCC. That comment says the A-profile defaults are already
correct; arm-none-eabi-gcc defaults to -mfloat-abi=soft, which disables the FPU
outright, so every VFP file fails with "selected processor does not support
'vmrs r1,FPSCR' in ARM mode". -mfloat-abi=hard alone is the fix and is the
right one, because it selects the core's own default FPU rather than naming a
-d16 one -- which is the trap the clang script warns about, since the
A-profile paths save D16-D31. Cortex-R4 is the exception in both scripts and
for the same reason: its FPU is an option rather than part of the core, so an
explicit -mfpu is required. Every value was measured against 14.3.rel1.
Stage 4 *unsets* TOOLCHAIN rather than setting it. The example build scripts
already default to GNU, and a stray TOOLCHAIN=atfe from a developer's shell
would otherwise make this stage silently check the other compiler. It cleans
the example directories on both sides, because the success test is the
existence of sample_threadx.out rather than the driver's exit status, and a
stale image from a previous toolchain would report success. Failure logs are
printed unfiltered: a missing tool says "command not found", and GNU ld's
undefined-symbol lines carry no "error:" at all.
Every skip is printed by name with a reason, per the house rule check_clang.sh
states three times -- a port simply absent from the count reads as covered.
This script also says outright that arm9 and arm11 are Arm and are skipped for
having no PORT_TARGET entry, which the clang script's "not Arm" wording glosses.
Verified on this tree with arm-gnu-toolchain 14.3.rel1: all five stages green,
every count identical to check_clang.sh's on the same tree -- 840, 469, 185x9,
42, 5 -- in 4m28s.
The failure paths were tested, not assumed. A deliberately broken .S in a
module port is reported by name and line in stages 1 and 2 and exits 1, in
--quiet mode as well. Reverting the AArch64 _init/_fini fix on one port only
gives "FAIL: cortex_a53: example build produced no image", 41 of 42, and exit
1 -- and the log tail it prints contains no "error:" anywhere, which is why it
is not filtered. A missing or wrong toolchain path exits 1 naming which triple
was not found.
RISC-V is deliberately out of scope for this first version: both ports
assemble 8 of 8 with the project's own cmake flags, but adding them widens the
toolchain download and the review surface for a family that is not regressing.
No regression test accompanies this. The script is the test, it exercises no
runtime behaviour, and its own failure paths are exercised above.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Ran the GCC port check in CI, on dev as well as master
scripts/check_gcc.sh with nothing invoking it would be a script nobody runs.
This adds the workflow, modelled on clang_check.yml, and fixes a trigger gap in
that file at the same time.
One job, two cache steps. Arm ships AArch32 and AArch64 as separate downloads
and the script needs both, so two caches keep the checks list short and let a
single invocation see both compilers. The AArch32 cache path and key match
cortex_m's exactly, so the two workflows share one entry rather than each
holding its own copy of the same archive -- noted in a comment, because the
only symptom of breaking that is a slower run.
Both triggers name dev. A workflow that triggers only on master gates no pull
request anybody opens; that is the defect ports_arch_check.yml carries a
comment about, and it cost cortex_m three months of failing in seven seconds
unnoticed. push is included as well as pull_request so dev's own history has a
baseline and a bad squash-merge is caught rather than waiting for the next PR.
The checksum suffix is .sha256asc and it is not interchangeable with .sha256.
Arm publishes both for this release, and verified 26 Aug 2026, the .sha256 file
for arm-none-eabi contains a 32-character MD5 rather than a SHA-256, so
sha256sum -c on it fails with "no properly formatted checksum lines found".
.sha256asc is a plain sha256sum-format line for both triples. The plan warned
that this suffix had changed between releases; the sharper truth is that both
suffixes exist simultaneously and one of them is not a SHA-256 at all. Recorded
in a comment beside the step.
Verified before writing them in rather than copied: both archive URLs and both
checksum URLs resolve, the archives are xz, the checksum files are
sha256sum-format for .sha256asc, and the AArch64 archive extracts to
arm-gnu-toolchain-14.3.rel1-x86_64-aarch64-none-elf/bin/aarch64-none-elf-gcc,
which is the path the workflow builds.
The paths: lists are duplicated between push and pull_request rather than
shared through a YAML anchor, deliberately: GitHub Actions' parser does not
dependably honour anchors and the failure mode is the workflow refusing to
parse, which is the cortex_m failure again. Ten duplicated lines are cheaper.
clang_check.yml's paths: list was missing CMakeLists.txt, cmake/ and
common_smp/, so that check did not run when files it reads changed -- the
ports_smp example builds compile common_smp/src and its CMake stage reads the
toolchain file and the top-level project. Both lists are now identical apart
from each file's own name, and both say so.
cortex_m is kept rather than deleted, against the plan's recommendation. It
builds four ports *through CMake*, and that is the only thing exercising
cmake/cortex_m*.cmake and the top-level CMakeLists for the M profile; this
script's CMake stage covers cortex_r52 only. The overlap is the assembly and
the C sources, not the build system, so deleting it would lose coverage rather
than remove a duplicate. Said so in the workflow header.
The script is passed explicit toolchain paths rather than left to find the
drivers on PATH, so nothing about the runner image can decide which compiler
runs, and it prints both versions it resolved.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The coverage summary reported a percentage and could not fail. Coverage
could fall from 99.97% to anything at all and every check stayed green,
against an AGENTS.md that asks for 100% test coverage -- a stated
requirement measured with a gauge that had no failure mode.
CodeCoverageSummary already takes thresholds and fail_below_min; neither
was set. Both are now, through a new coverage_thresholds input on the
template, because the two suites do not sit at the same figure: ThreadX
99, SMP 98.
Three things were probed against the pinned action on a runner before
picking those numbers, using the real merged.xml files from the dev push
run of #666.
The floor compares the line rate and nothing else. That mattered because
branch coverage is around 78% in both suites while line coverage is
98.8-100%, so a floor aimed at the line figure would have been an
immediate red wall had it tested branches or the lower of the two. The
ThreadX report at 100.00% lines and 77.67% branches clears a floor of 99.
The thresholds are whole numbers. '99.9 100' -- the value this was meant
to be -- is rejected with 'System.ArgumentException - Threshold parameter
set incorrectly.', and the step fails whether or not fail_below_min is
set. So the choice is 99 or 100 with nothing between.
100 would fail on a race. tx_thread_system_resume.c:529 is reached by
timing rather than by construction and flaps between runs of the same
green tree, which is why #666 left it; 4502/4503 fails a floor of 100 and
clears one of 99. A coverage gate that goes red on a coin toss is how
coverage gates get switched off.
SMP is 5114/5178 lines, 98.76%, with 64 uncovered lines across 11 files
of common_smp/src -- #666 closed the equivalent gaps in common/src only.
A shared floor of 99 would have failed that job on every run while
ThreadX passed.
One limit is recorded in the file rather than fixed: an empty report
reads as 100%. gcovr writes line-rate="1.0" beside lines-valid="0" when
it finds no data, and the action prints 'Line Rate = 100% (0 / 0)' and
passes any floor. The check for that is the emptiness assertion #664 put
in each suite's coverage.sh, not this one.
Also corrected two stale filenames in the deploy job's comment: since
#665 each coverage artifact carries merged.xml, not
default_build_coverage.xml. Verified on the runner.
Assisted-by: Claude Opus 5 <noreply@anthropic.com>
Only default_build_coverage carried -fprofile-arcs, because the gate was the
build type and it is the only one of five whose name ends in _coverage. The
other four build and run all their tests and their coverage was discarded. That
is not redundancy thrown away: each configuration selects a different set of TX_
feature macros, so the code the other four compile is absent from the
denominator rather than uncovered in it.
TX_COVERAGE instruments a build regardless of its name, defaulting to OFF so a
single configuration built by hand behaves as before. coverage.sh gains a
--merge mode that unions the per-configuration JSON tracefiles, and
cmake_bootstrap.sh runs it after the test loop so a local run produces the same
merged report CI reads. The template sets TX_COVERAGE for build and test, and
coverage_name moves to the merged report.
Measured on the ThreadX suite, all 480 tests passing:
default_build_coverage 3827 valid 3827 covered
disable_notify_callbacks_build 3767 3766
stack_checking_build 3857 3856
stack_checking_rand_fill_build 3862 3861
trace_build 4123 4108
merged 4503 4487
The denominator grows by 676 lines, 17.7%, and the figure moves from 99.97% to
99.64%. The second one is honest, and the drop is the point rather than a
regression: the denominator now includes code the old report never counted. The
union also contains a file the old report did not contain at all --
tx_thread_stack_error_handler.c compiles only under TX_ENABLE_STACK_CHECKING, so
it was not listed at 0%, it was simply absent. 177 files becomes 178.
Coverage collection moved out of test() and now runs after the test loop, one
configuration at a time. gcov writes its intermediate gcov files into the
directory gcovr is rooted at, and coverage.sh roots every configuration at the
repository root so filenames come out repo-relative. Five concurrent gcovr
processes therefore share one scratch directory and delete each other's output:
the first full run of this change passed all 480 tests and produced no report
for three of the five configurations. Measured both ways -- two gcovr rooted at
the repository root fail concurrently and succeed in sequence. CI would not
have caught it, because test_tx.sh sets CTEST_PARALLEL_LEVEL=1 and takes the
serial branch.
Per-configuration output moved under coverage_report/per_configuration/ and is
excluded from the Pages artifact. The deploy job merges the ThreadX and SMP
artifacts into one tree and every configuration directory has the same name in
both, so left at the top level one suite's would overwrite the other's on the
published site.
On the SMP suite, an earlier run of this change saw trace_build fail
threadx_smp_time_slice_test and then hang, which raised the question of whether
-fprofile-arcs perturbs a timing-sensitive test. It does not. Sixteen runs
settle it, and the decisive one is that threadx_smp_time_slice_test failed
ERROR #31 -- twice in a row under --repeat until-pass:2 -- on an uninstrumented
build, in the exact shape CI runs, while three instrumented runs of that shape
passed 5 of 5. In the CI shape, CTEST_PARALLEL_LEVEL=1 run.sh test all:
TX_COVERAGE=OFF 3 runs 2 green, one ERROR #31 310 s
TX_COVERAGE=ON 3 runs 3 green, 5/5 each 325-329 s
So the test is a pre-existing flake on dev and instrumenting all five costs
about 5% of the suite's wall clock. Separately, and also in both instrumented
and uninstrumented builds, run.sh's parallel branch -- what a developer gets
typing run.sh test all with no CTEST_PARALLEL_LEVEL -- hangs under its own load,
four times in twelve runs. Several SMP tests create 1024 ThreadX threads by
construction and the Linux port backs each with a pthread, so five
configurations at once put on the order of 5000 threads on the machine. CI sets
CTEST_PARALLEL_LEVEL=1 and does not take that branch.
Assisted-by: Claude Opus 5 <noreply@anthropic.com>
The action references were pinned to commit SHAs in #660, and a SHA pin with
nothing moving it is worse than a floating tag -- it holds CI on whatever was
current the day it was written. That is exactly how actions/cache@v1 stayed in
ci_cortex_m.yml until GitHub began auto-failing every request that used it.
The drift measured before that catch-up: download-artifact four majors behind,
checkout and upload-artifact three each, cache and upload-pages-artifact two,
with nothing ever reporting it. This closes the loop, and the reference to
.github/dependabot.yml that #660 left in each workflow's pinning comment.
Weekly, github-actions only. Patch and minor are grouped into one pull request
because they are the routine traffic and a queue reviewed one item at a time is
a queue that gets ignored. Majors stay ungrouped, one each, because every
breaking change this repository has met in an action has been a major.
Two choices worth stating rather than leaving to be rediscovered.
target-branch is dev. Dependabot reads this file from the default branch, which
is master, but master is deliberately kept behind dev and pull requests belong
where the regression suites gate them. The consequence is that landing this on
dev arms it without firing it: nothing happens until a release merge carries
the file to master. Setting target-branch also opts out of Dependabot security
updates, which only run against the default branch -- a small cost for this
ecosystem, since an action advisory arrives as an ordinary bump on the weekly
run, but a real one.
The pull-request limit is raised from the default five to ten. Nine actions are
in use, and five would hold majors back with nothing saying that it had.
No other ecosystem is configured, deliberately: external dependencies are
forbidden, there are no submodules, and the one pinned tool -- gcovr in
scripts/install.sh -- lives in a shell script no ecosystem can parse, so that
pin keeps moving by hand.
No sibling eclipse-threadx repository has a Dependabot configuration, so this
sets the pattern rather than following one. The dependencies label it uses
already exists here.
Assisted-by: Claude Opus 5 <noreply@anthropic.com>
Node 20 is removed from the GitHub runners on 16 September 2026. Every run
in this repository currently emits the deprecation warning for it, naming
actions/checkout, actions/configure-pages, actions/upload-artifact,
LouisBrunner/checks-action and marocchino/sticky-pull-request-comment among
others. After that date those actions stop working rather than warning, so
this is a deadline and not housekeeping.
Every action is now referenced by a 40-character commit SHA with the version
in a trailing comment. A tag can be repointed at any commit; a SHA cannot, so
this is what makes "which code ran in our CI" answerable from the repository
rather than from whatever the tag meant at the time. The versions were behind
by as much as four majors -- download-artifact was on v4.3.0 against v8.0.1 --
because nothing in this repository has ever reported that an action moved.
Compatibility was checked against each new action.yml rather than assumed,
for every input this repository actually passes:
checkout submodules is unchanged
cache path and key are unchanged
upload-artifact name, path and retention-days are unchanged
download-artifact pattern, merge-multiple and path are unchanged
configure-pages takes no input here, and none became required
deploy-pages still exposes page_url, which the job reads
upload-pages-art. path is unchanged
checks-action token, name, conclusion, output and
output_text_description_file all survive v2 to v3
sticky-comment header and path survive v2 to v3, and the new
GITHUB_TOKEN input defaults to github.token, which is
what v2 used implicitly
delete-artifact name survives v5 to v6, and useGlob still defaults to
true, so the coverage_report-* glob from #655 still
matches
CodeCoverageSummary already current at v1.3.0; pinned, not moved
The artifact pair moves together, as it must. The round trip was verified on
a runner before this commit: upload-artifact v7 to download-artifact v8,
through the pattern and merge-multiple selection #655 introduced, filtered 4
artifacts to 2 and produced exactly the tree the deploy expects.
Two behaviour changes worth knowing. download-artifact v8 adds a
digest-mismatch input defaulting to error, so a corrupted artifact now fails
the job instead of passing through -- the right default, but a change.
upload-artifact v6 and above require a runner of at least 2.327.1, which the
hosted runners satisfy and a self-hosted runner would need checking for.
Assisted-by: Claude Opus 5 <noreply@anthropic.com>
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
Two defects, and the second hid the first.
This workflow triggered on master only, for both push and pull_request,
while dev is the integration branch. So it gated no pull request that
anybody opened. That is the same defect ports_arch_check.yml carries a
comment about, where it cost eight months of ports drifting from ports_arch
unnoticed, and regression_test.yml has it too.
And it had not compiled anything since at least 2026-06-08. Every run since
then failed in six to eight seconds at "Prepare all required actions",
before checkout, because GitHub automatically fails any request that uses
actions/cache@v1. The last run of any kind was 2026-06-30. A workflow that
fails in seven seconds is normally noticed within the day; this one was not,
because of the first defect. The two together meant the project's only job
that cross-compiles a port with GCC had been reporting nothing at all.
The toolchain now follows clang_check.yml rather than third party actions:
a pinned release fetched directly from Arm, verified against the published
sha256asc, and cached with actions/cache@v4. That also moves the compiler
off 9-2019-q4, a 2019 release, onto a version matching the GCC 14 default
this project states.
The ninja install is guarded on ninja being absent rather than run
unconditionally, because scripts/install.sh already carries a long comment
about apt-get update stalling for over two hours and taking a whole
regression run with it.
fail-fast is off so that one port failing still reports the other three.
Verified before committing, with the pinned 14.3.rel1 toolchain: the
download and sha256sum -c sequence in the install step was run as written,
the archive extracts to the directory the PATH step expects, and all four
ports configure and build clean with zero warnings.
This covers four ports of the forty under ports/ that have a gnu directory.
Widening it to every Arm gnu port is separate work.
Assisted-by: Claude Opus 5 <noreply@anthropic.com>
A failing test threw away coverage that had already been collected, and the
run whose behaviour changed is exactly the run whose coverage is worth
reading. Measured on the failing run of 2026-08-18: it uploaded test_reports
for all three suites and no coverage_report artifact at all.
Two causes, and the workflow one is the smaller of them.
cmake_bootstrap.sh runs under set -e, so a failing ctest aborted test()
before ./coverage.sh was reached. The gcda files exist by that point, so
nothing was missing except the step that reads them. ctest's status is now
captured and returned at the end, and the summary grep is allowed to fail
rather than being the thing that stops the coverage behind it.
The serial branch of the test dispatch collected no status either, so under
set -e the first failing configuration stopped the remaining four from being
tested at all -- and their coverage from being collected. That was cheap
while the suites ran in parallel, because the parallel branch already
collects exit codes from its background jobs. Moving to serial execution in
#643 quietly made one failure cost the other four configurations. The serial
branch now collects status the same way the parallel branch does.
With those fixed the report exists, so the workflow steps that publish it no
longer skip on failure. They are guarded with !cancelled() rather than
always(), so a cancelled run still stops promptly, which is the idiom
deploy_code_coverage already uses. The ${{ }} wrapping is required and not
decoration: a bare ! opens a YAML tag, and the file will not parse without
it.
Verified locally by replacing one test binary with a stub that exits 1:
before the failing run of 2026-08-18 produced no coverage_report artifact
after run.sh test default_build_coverage exits 8, and produces
coverage_report/default_build_coverage.xml with 177 files and
3804 of 3827 lines
after run.sh test all exits 8, and all five configurations run rather
than stopping at the first
The failure still fails. Only the reporting around it changed.
Assisted-by: Claude Opus 5 <noreply@anthropic.com>
The download step in deploy_code_coverage asked for the artifact named
${{ steps.artifact.outputs.coverage_report }}. That output is set by the
"Coverage Report name" step of run_tests, which is a different job, and the
steps context does not cross jobs. So the expression evaluated to the empty
string and the action took its documented path for an unspecified name:
No input name, artifact-ids or pattern filtered specified,
downloading all artifacts
Total of 4 artifact(s) downloaded
The four are the two coverage reports and the two test_reports bundles of
JUnit XML, each extracted into a directory named after the artifact. The
next step uploads the lot to Pages, so the published site has carried the
test reports alongside the coverage, one directory deeper than intended,
under a path containing a run timestamp that changed on every publish. Any
link to a coverage report broke the next time one was published.
Selecting by pattern with merge-multiple fixes both halves: the pattern
excludes the test_reports bundles, and merging puts the contents of the two
coverage artifacts directly into coverage_report rather than under a
directory named for each. Each artifact holds one directory named for its
suite, renamed from default_build_coverage by "Prepare Coverage GitHub
Pages", so the result is the two suite directories the deploy expects and
the timestamped artifact name no longer appears in the published path.
Verified on a runner rather than reasoned about, with an isolated workflow
that uploads artifacts shaped like the real ones and downloads them both
ways:
OLD coverage_report/coverage_report-<epoch>-ThreadX/ThreadX/index.html
coverage_report/coverage_report-<epoch>-ThreadX/default_build_coverage.xml
coverage_report/coverage_report-<epoch>-SMP/SMP/index.html
coverage_report/coverage_report-<epoch>-SMP/default_build_coverage.xml
coverage_report/test_reports SMP/results.xml
coverage_report/test_reports ThreadX/results.xml
NEW coverage_report/ThreadX/index.html
coverage_report/SMP/index.html
coverage_report/default_build_coverage.xml
Both artifacts carry a default_build_coverage.xml and the merge means one
overwrites the other, which the run above also shows. That file is consumed
by CodeCoverageSummary back in run_tests and is not read here, so it is
untidy rather than wrong, and it is called out in a comment.
The delete step is fixed in the same place and for a related reason. The
artifacts are named coverage_report-<epoch>, useGlob defaults to true in
this action, and as a glob "coverage_report" matches only the literal
string. It has been deleting nothing, without failing, and retention-days: 1
on the upload is what has actually been clearing these up.
Assisted-by: Claude Opus 5 <noreply@anthropic.com>
Running the regression suites on dev (#652) was meant to test the branch the
pull requests target. It changed what gets published as well, which was not
intended and does not work: the first push to dev after that merge failed
with
Branch "dev" is not allowed to deploy to github-pages due to
environment protection rules.
All three suites passed in that run -- tx, smp and freertos. The only
failure was deploy / deploy_code_coverage, rejected before it ran, because
the github-pages environment restricts deployments to master.
The guard goes here rather than in the environment settings, because the
environment rule is doing its job. Which branch the published coverage
report describes is a deliberate decision, and moving it from master to dev
is a change worth making on purpose rather than as a side effect of a
trigger fix. Doing so needs the environment setting relaxed as well as this
line removed.
The per-suite deploy_code_coverage jobs need no guard: tx, smp and freertos
all pass skip_deploy: true, and regression_template.yml already restricts
that job to push and workflow_dispatch. Only the deploy job, which is the
one that publishes, was reaching the environment.
Assisted-by: Claude Opus 5 <noreply@anthropic.com>
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The ThreadX, SMP and FreeRTOS-compatibility suites triggered on master only,
for both push and pull_request. dev is the integration branch, so these
suites gated no pull request that anybody opened: the last dev run of any
kind was a manual workflow_dispatch on 2026-08-18.
This is the same defect ports_arch_check.yml already carries a comment
about, where it cost eight months of ports drifting from ports_arch
unnoticed. ci_cortex_m.yml has it too and is handled separately.
That 2026-08-18 run failed, which is the reason to check before switching
this on rather than after. Two tests failed: threadx_timer_simple_test in
the ThreadX suite, with ERROR #28, and threadx_thread_priority_change in
the SMP suite, with a timeout. Both were fixed two days later -- the first
by running the suites one test at a time (#643), which is what a timer test
failing only under parallel load wants, and the second by #647 by name.
Verified before this commit rather than assumed: both suites were re-run on
this tree, and all 1030 tests pass across all ten build configurations, the
ThreadX suite in 34 to 64 seconds per configuration and the SMP suite in 61
to 63. The 2026-08-18 run took 36m19s, of which a single test that has since
been given a budget accounted for 439 seconds.
No paths filter is added deliberately. The suites build the linux port, so a
filter would have to enumerate what cannot affect them, and the failure mode
of getting that list wrong is a regression that merges because the filter
excluded the file that caused it.
The deploy job needs no guard: regression_template.yml already restricts
deploy_code_coverage to push and workflow_dispatch, and restricts the
coverage PR comment to pull requests from the repository itself, so neither
fires for a pull request from a fork.
Assisted-by: Claude Opus 5 <noreply@anthropic.com>
The Install softwares step runs apt-get update, an apt-get install and a pip
install, and it has stalled twice: 55 minutes on the SMP job of one run, and
more than two hours on the ThreadX job of the next, against a normal 27 to 152
seconds across every other run measured. In the second case the tests never
started at all.
No step in this template had a timeout, so a stall runs until the six hour job
limit. That turns a transient apt or PyPI problem into a lost run, and it hides
what happened: the job simply sits there, and the failure that eventually gets
reported says nothing about which step was stuck.
Bound the three steps that do real work. Ten minutes for the install, against a
normal worst case of 152 seconds. Fifteen for the build, which has run between 4
and 54 seconds. Sixty for the test step, which is the only one whose length
depends on the suites themselves; the longest observed is 37 minutes, and that
was with a wait in one test that has since been bounded.
A step that trips its timeout fails and names itself, which is the point. None
of these numbers is tight enough to trip on work that is merely slow.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
The gnu ports are only ever built with GNU tooling, and GNU as accepts several
non-canonical forms that LLVM's assembler rejects. Nothing noticed, because
nothing built them with anything else. This matters beyond clang itself: Arm
Toolchain for Embedded is LLVM based and is the successor to Arm Compiler 6, so
these are the code paths ac6 users move onto.
Seven files needed changing, none of which alters the emitted code:
LDREX and STREX take no offset in A32 state; the #imm form is Thumb-2 only. GNU
as drops the redundant zero, LLVM rejects it. Removed from the Cortex-A5, A7 and
A9 SMP protect routines.
ARMv8-M Baseline has no flag-preserving MOV immediate, so GNU as already emits
MOVS. Writing MOVS in the two Cortex-M23 sources says what the assembler was
doing anyway. One of them sits in a branch only compiled for the single mode
secure configurations, which is why it had never surfaced.
The Cortex-M0 schedule routine wrote LDR r0, =#0x10000000 with a stray hash,
which its own sibling file already wrote correctly.
The Cortex-M0 system return routine selected the numbered subsection .text 32,
which makes LLVM place the constant pool beyond the range a Thumb-1 PC relative
load can reach. Plain .text fixes it and GNU accepts either form. The reason is
recorded in the file, since 32 files pair a numbered subsection with a literal
pool load and the rest only escape because Thumb-2 and A32 have far more range.
Add scripts/check_clang.sh, which assembles every Arm gnu port source and
compiles the common C sources for one core per architecture profile, and a
clang_check workflow that installs Arm Toolchain for Embedded and runs it. The
toolchain version is pinned and checksum verified, for the same reason the
runner image is pinned.
The port directory to target mapping in that script is explicit rather than
prefix matched. Prefix matching is what makes cortex_a5 also match cortex_a53
and cortex_a55, which are AArch64, and assembling those as ARM32 produces
hundreds of misleading errors; that mistake cost real time while measuring this,
so the reason is recorded next to the table.
Verified with Arm Toolchain for Embedded 22.1.0: 711 of 711 assembly sources
assemble and 185 of 185 C sources compile for all eight profiles, against 6
assembly failures before the change. arm-none-eabi-gcc still assembles all 315
ARM32 sources, so nothing regressed for GNU. The check was confirmed to fail
when any one of the fixes is reverted.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
The ARMv7-A and ARMv8-A ports are generated by update.ps1, which needs
PowerShell, so the cortex-a job in ports_arch_check ran on a Windows image and
nobody could reproduce it locally on Linux. Add update.sh beside each
update.ps1, with the same cores, compilers, copy sets and patches, and move the
job to the same Linux image as everything else.
The bash scripts were checked against the PowerShell ones by comparing what
each reports as drifted. They agree exactly on the 63 files the Windows job
last reported, and differ on 12 more, which turn out to be a defect in
update.ps1 rather than in the port. Its two .cproject patterns are written as
'value=`"cortex-a7`"' with backticks that survive into the pattern, so that
replacement has never matched, while the neighbouring Cortex-A7.NoFPU pattern
has no backticks and always worked. The result is that the AC6 example builds
for the A5, A8, A9, A12, A15 and A17 cores name cortex-a7 as their CPU while
their FPU string is correct. The bash scripts do what the PowerShell ones
intended, so regenerating corrects those twelve files.
Restore ports_arch as the source for the rest. The implementation of
_tx_thread_smp_time_get from #555 was applied to the twenty four generated SMP
ports and never to ports_arch, which still held MOV x0, #0 with a FIXME
comment, so regenerating would have replaced a working generic timer read with
a stub. That implementation now lives in the source. The remaining differences
are cosmetic and resolve in favour of the source: a trailing blank line in 38
copies of tx_thread_schedule.S and comment spacing in one tx_port.h.
Note that the Cortex-A VFP fix is already present in ports_arch and was never
at risk, contrary to what the description of the port consistency checks change
said before this was measured.
Extend scripts/check_ports.sh to run the A profile generators too, and make it
fail when a generator fails or is missing rather than reporting a clean tree,
which would have been a false pass.
Pin every workflow to ubuntu-24.04. ubuntu-latest already resolves to that
image, so nothing changes today, but a future migration becomes a deliberate
commit rather than something that happens underneath the -m32 builds.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
Three defects reached the repository through the port trees recently, and each
of them is mechanically detectable without a cross compiler. Add
scripts/check_ports.sh, which looks for exactly those three, and give CI and
the release process the same command a contributor can run locally.
The generated Cortex-M ports must be reproducible from ports_arch. Fixes were
applied to the generated copies instead of the source for eight months, and the
next run of the copy scripts would have reverted them.
Preprocessor directives must balance. A fix left the Cortex-M85 IAR tx_port.h
with one more #endif than #if, so that header could not compile.
No port header may carry a statement outside a function body. A fix left a
second, headerless copy of a function body in the Cortex-M4 AC6 tx_port.h,
which is issue 569. The check tracks brace depth while skipping preprocessor
lines, multi-line macro bodies and comments, and reports assignments,
dereferences and control statements that land at file scope. Headers under
example_build are excluded, since those trees vendor third party SDK code.
A fourth section reports, without failing the run, on port families that have
no copy script and so cannot be checked for reproducibility. It currently
observes that the Cortex-M0 ac5, ac6 and keil ports lack the barriers their gnu
and iar siblings have.
ports_arch_check now calls the script rather than inlining a copy and diff, so
CI and the command line check the same things by the same definition, and the
workflow now triggers on pull requests to dev as well as master. Triggering on
master alone is why the drift went unseen. prepare_release.sh runs the checks
before it branches or rewrites anything, and stops if they fail, with
SKIP_PORT_CHECKS=1 as the escape hatch.
Each check was verified by reintroducing the defect it exists to catch and
confirming that the script fails, then confirming it passes on a clean tree.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Fixed the resource leaks on the xQueueCreate error paths
xQueueCreate() allocated the queue descriptor and its backing memory, then
created two ThreadX semaphores, and returned NULL on either semaphore failure
without releasing anything. Since no handle reached the caller, vQueueDelete()
could not be used to recover, so both allocations were lost. A failure on the
second semaphore additionally abandoned the read semaphore it had already
created, leaving a live ThreadX control block inside freed memory.
Release the backing memory and the descriptor on both paths, and delete the
read semaphore before returning when the write semaphore cannot be created.
This is the teardown order vQueueDelete() already uses, and it matches the
cleanup xTaskCreate() performs on its own error paths.
Verified with a fault injection harness that intercepts the ThreadX byte pool
and semaphore entry points to force tx_semaphore_create() to fail on a chosen
call. On a read semaphore failure the layer previously performed 2 allocations
and 0 releases, and on a write semaphore failure 2 allocations, 0 releases and
0 semaphore deletions. It now performs 2 releases in both cases and deletes
the read semaphore in the second, with the byte pool restored to its prior
state.
Fixes https://github.com/eclipse-threadx/threadx/issues/570
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Added a regression suite for the FreeRTOS compatibility layer
The compatibility layer had no tests in this repository, which is awkward for
its creation functions in particular. Each of them takes one or two byte pool
allocations for its bookkeeping and then creates ThreadX kernel objects, and
each returns NULL when a kernel object cannot be created. The caller is left
without a handle, so it cannot call the matching delete function, and anything
the layer failed to release is gone until the system restarts. A leaking
version and a correct version are indistinguishable from the outside, which is
how the leak in issue 570 went unnoticed.
Add a suite that counts what the layer takes and gives back. A test asks the
harness to fail a chosen kernel creation call, then checks the number of byte
pool allocations, releases, object creations and object deletions performed.
The ThreadX entry points are intercepted with the linker's --wrap so that
tx_freertos.c is compiled exactly as it ships, with no test hooks in it. Note
that tx_api.h maps the public API onto the error checking entry points, so the
_txe_ symbols are the ones wrapped. Coverage is the creation and teardown paths
of queues, tasks, semaphores, mutexes, event groups and timers, including a
regression test for the two paths fixed for issue 570.
The suite follows the layout of the existing ThreadX and SMP suites, is
registered with ctest, and runs in CI through the shared regression template.
It is built 32 bit because the Linux port defines ULONG as unsigned int on
x86_64 while the layer passes pointers through ULONG arguments, so a 64 bit
build truncates them. It is Linux only because --wrap has no MSVC equivalent,
and the CMake configuration says so rather than failing at link time.
Validated by building the suite against the layer as it stands before the
issue 570 fix, where the two expected checks fail with the leaked counts, and
against the fixed layer, where all three tests pass.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
risc-v: refactor, consolidate, and fix RV32/RV64 ports
Consolidates the RISC-V 32-bit and 64-bit GNU/Clang port sources, fixes two
pre-existing assembly bugs discovered during testing, and hardens the build
infrastructure for both the regression suite and the CORE-V MCU example.
--- Port consolidation (RV32 GNU + Clang) ---
- Delete ports/risc-v32/clang/src/ (8 .S files had no Clang-specific
directives; diverged from GNU only due to missing bug fixes). The Clang
port CMakeLists.txt now compiles from ../gnu/src/.
- Change .global -> .weak for _tx_initialize_low_level in gnu/src/ to allow
BSP-level override without a linker conflict (adopted from Clang port).
- Create ports/risc-v32/common/tx_port_riscv32_common.h with all definitions
shared between GNU and Clang ports. Reduce both tx_port.h files to thin
wrappers.
- Add a prominent comment in risc-v64/gnu/inc/tx_port.h explaining why
LONG/ULONG are intentionally 32-bit on RV64 (ThreadX ABI requirement,
mirrors win64/MSVC LLP64).
--- Shared CMake helper ---
- Add cmake/threadx_riscv_port.cmake with threadx_add_riscv_port(). All
three port CMakeLists.txt files are reduced to ~8 lines each. Include path
is relative to CMAKE_CURRENT_LIST_DIR so the helper works whether ports
are built standalone or as a subdirectory of the test framework.
--- Shared example-build drivers ---
- Create canonical driver files under ports/risc-v_common/:
inc/csr.h (uintptr_t-based; portable RV32 + RV64)
example_build/plic/ (plic.c, plic.h)
example_build/uart/ (uart_qemu_ns16550.c/h; static inline putc_nolock)
example_build/trap/ (trap_qemu.c; XLEN-portable mcause constants)
- Replace per-example copies with symlinks in all qemu_virt and cva6_ariane
example directories.
- Fix OS_IS_INTERRUPT typo (was OS_IS_INTERUPT) in shared trap_qemu.c.
- Gate print_hex() behind TX_RISCV_TRAP_DEBUG.
--- Bug fixes in RV32 assembly ---
tx_thread_schedule.S:
- Solicited-return FP path: reload t0 from the mepc stack slot before
csrw mepc, t0. After the FP restore block, t0 held the fcsr value (0 for
new threads), which caused mepc = 0 and an immediate instruction-address
fault on the first context switch.
- Same path: reload t0 from the mstatus stack slot before csrw mstatus, t0
to avoid writing the stale fcsr value into mstatus.
tx_thread_system_return.S:
- FP callee-saved registers were saved unconditionally before the mstatus.FS
check, causing an illegal instruction trap (mcause=0x2) when a thread with
FS=Off (lazy FPU, thread has never used FP) voluntarily yielded.
- Apply the same FS guard pattern used in tx_thread_context_save.S: read
mstatus first, isolate FS[1:0], and skip fsw/fsd if FS == Off.
Both bugs were pre-existing on origin/dev and are unrelated to the
consolidation changes.
--- RV64 64-bit pointer compatibility ---
- Add TX_TIMER_INTERNAL_EXTENSION, TX_THREAD_CREATE_TIMEOUT_SETUP, and
TX_THREAD_TIMEOUT_POINTER_SETUP to risc-v64/gnu/inc/tx_port.h to store the
thread timeout pointer in a VOID
* extension field rather than truncating it
into a 32-bit ULONG. Mirrors the win64 port pattern.
- Define TX_TIMER_EXTENSION_PTR_DEFINED as a portable sentinel.
- Update threadx_thread_basic_execution_test.c guard from #if defined(_WIN64)
to #if defined(_WIN64) || defined(TX_TIMER_EXTENSION_PTR_DEFINED).
- Disable -Wconversion for the RV64 test build: ULONG = unsigned int (32-bit)
is intentional for ThreadX ABI but triggers spurious warnings when sizeof()
(8 bytes on RV64) appears in arithmetic with ULONG in common/src/.
--- Regression suite cmake fixes ---
test/tx/cmake/riscv/regression/CMakeLists.txt:
- Build testcontrol_weak_defaults.c as a separate OBJECT library and include
it in every test executable via $<TARGET_OBJECTS:>. GNU ld does not extract
objects from a static archive to satisfy weak symbols, so bundling it in
test_utility was insufficient for the standalone
threadx_initialize_kernel_setup_test.
test/tx/cmake/regression/CMakeLists.txt,
test/smp/cmake/regression/CMakeLists.txt:
- Same fix applied to the Linux and SMP regression builds. The symbols
abort_all_threads_suspended_on_mutex, suspend_lowest_priority, and
abort_and_resume_byte_allocating_thread were introduced by the win64 merge
and left the standalone test unlinkable.
--- CORE-V MCU toolchain and build fixes ---
cmake/riscv64-gcc-rv32imc.cmake:
- Resolve riscv64-unknown-elf-gcc via PATH so the riscv-collab toolchain in
/opt/riscv/bin is preferred when it appears first.
ports/risc-v32/gnu/example_build/core_v_mcu/bsp/clz.c (new):
- The riscv-collab toolchain is built without rv32 multilib, so its libgcc
does not define __clzsi2 (the helper emitted for __builtin_clz() in fll.c).
Add a weak __clzsi2 fallback so the build is self-contained with any
riscv64-unknown-elf toolchain. The weak attribute yields to a
libgcc-provided strong symbol when the Ubuntu multilib package is used.
core_v_mcu/CMakeLists.txt:
- Add bsp/clz.c to sources.
- Reference CMAKE_TOOLCHAIN_FILE via message(STATUS) to suppress the false-
positive "Manually-specified variables were not used by the project" CMake
warning and to show the active toolchain at configure time.
--- Housekeeping ---
- Rename azrtos_test_* -> threadx_test_* (eliminate Azure RTOS branding).
- Add RV64 QEMU CI test script:
ports/risc-v64/gnu/example_build/qemu_virt/test/
threadx_test_tx_gnu_riscv64_qemu.py
- Normalize entry.s -> entry.S in all 4 example directories.
- .gitignore: exclude build_m7/ and .codex local artifacts.
- CI: comment out the riscv regression workflow job and remove it from the
deploy job's needs list (preserved in-place for easy re-enablement).
--- Verified ---
- 95/95 RV32 regression tests pass (QEMU virt)
- 95/95 RV64 regression tests pass (QEMU virt)
- All 5 Linux build configurations build cleanly (default_build_coverage,
disable_notify_callbacks_build, stack_checking_build,
stack_checking_rand_fill_build, trace_build)
- CORE-V MCU example_build links cleanly with /opt/riscv toolchain
Co-authored-by: Copilot 223556219+Copilot@users.noreply.github.com
* Test multiple code coverage pages
* Add affix to artifacts
* Test uploading code coverage as artifact
* Deploy GitHub pages at last for multiple jobs
* Test using unified upload pages
* Disable test cases to accelerate experiment
* Fix escape character $
* Revert "Test using unified upload pages"
This reverts commit 3668d9f672.
* Set destination for downloaded artifact
* Use a different artifact name
* Fix escape value
* Revert "Disable test cases to accelerate experiment"
This reverts commit 8468f17d02.
* Override duplicated github-pages in artifact
* Revert "Override duplicated github-pages in artifact"
This reverts commit 17a83aa97d.
* Delete Duplicate Code Coverage Artifact
* Convert ADO pipelines to GitHub actions
* Remove version in uses as not valid for local workflows
* Fix cmake path and add deploy url affix
* Add SMP build job
* Fix code coverage URL
* Add affix to titles of steps
* Remove ADO pipelines
* Add affix to titles of code coverage
* separate PR results for multiple jobs
* Revert "separate PR results for multiple jobs"
This reverts commit 6da13540fd.
* separate PR results for multiple jobs
* Unify ThreadX and SMP for ARMv8-A.
* Fix path in pipeline to check ports arch.
* Add ignore folders for ARM DS
* Generate ThreadX and SMP ports for ARMv8-A.
* Ignore untracked files for ports_arch check.
* Use arch instead of CPU to simplify the project management.