mirror of
https://github.com/eclipse-threadx/threadx.git
synced 2026-10-06 06:59:08 +08:00
cd6a2d90347945f2736a6afe25dcf745ffc31a6c
473
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
cd6a2d9034 |
Put the module manager test's thread on the kernel's created list, so the stand-in kernel matches the one the manager reads (#739)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The host test built a TX_THREAD, gave it an ID and a module instance, and left it off the kernel's created list. A created thread lives on that list, so the thread the test created was not one the kernel would recognise. The created list head and count are now defined for each object type, the thread the test creates goes onto the thread list the way thread create puts it there, and a successful delete takes it off again the way thread delete does. Only the thread list is populated; the others stay empty because nothing here creates an object of those types. No expectation changed. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
83a3e62c28 |
Added a Module Manager host test harness and released a module thread's kernel stack only when it has one (#738)
* Released a module thread's kernel stack only when it has one, so deleting a thread without a manager-allocated stack no longer asks for that memory back The thread delete dispatcher decided whether to release a kernel stack from the calling module's property flags. A user-mode module can be given the address of a thread that carries no kernel stack the manager allocated, and deleting it passed that thread's null stack pointer to _txm_module_manager_object_deallocate(). The pointer is now tested instead of the module's properties. Both thread create paths clear the whole control block before filling it in, so a thread with no manager-allocated kernel stack holds TX_NULL in the field, which makes the test exact rather than an inference from how the module was loaded. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> * Added a Module Manager host test harness and covered the user-mode thread kernel stack lifetime The Module Manager is not in the ThreadX library the test tree links, and its control blocks come from a module port rather than a base port, so nothing in the suite could reach it. The thread_transition directory already describes this technique as the one "the module manager tests in this tree use", but no such tests existed. This adds the directory that comment refers to: a port shim that replaces the three interrupt primitives with host equivalents that count, and a CMake target that compiles the manager sources under test directly against one module port's headers. The dispatch layer is one header of static functions and an unoptimised build emits all of them, so a test that includes it to reach one dispatcher would pull in references to every service the manager can dispatch. The header guards each dispatcher with its own TXM_<SERVICE>_CALL_NOT_USED macro, so the build reads that list out of the header and defines every guard except the ones the test needs. Reading it rather than writing it down means a service added later is excluded without anyone having to remember. The first test asserts the invariant a user-mode module thread has to hold: one logical thread costs the object pool two allocations, the control block the module asks for and the kernel stack the manager takes on its behalf, and deleting the thread must return the pool's available bytes and the module's allocation-list count to exactly what they were. It asserts an equality rather than a bound, because a bound would pass while one of the two was left behind on every cycle, and then runs enough cycles that a leak of one stack per cycle exhausts the pool several times over. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
dde43b8ab2 |
Replaced the stale system stack switch pseudo-code in the ARMv7-A ports with comments that describe what the code actually does (#735)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The context save, vectored context save and system return routines in the
ARMv7-A ports carried pseudo-code comments claiming that they saved the
thread stack pointer and then switched to _tx_thread_system_stack_ptr.
Neither of those things happens, and none of these ports references that
variable outside an unused IMPORT in their example builds.
On ARMv7-A each processor mode has its own banked stack pointer. The IRQ
handler branches to _tx_thread_context_save while still in IRQ mode, so the
core's banked IRQ stack already serves as the system stack, and the thread
stack pointer is stored in the control block by _tx_thread_context_restore,
and only when the interrupt results in preemption. The scheduler runs on the
banked SVC mode stack that the startup code sets up. There is nothing for a
software stack switch to do.
The comments were therefore misleading rather than merely redundant, and had
led at least one user to try to restore the code they described. They are now
replaced by a description of the actual mechanism.
The AArch64 SMP ports keep their comments unchanged, because ARMv8-A does not
bank a stack pointer per processor mode and those ports do reload
_tx_thread_system_stack_ptr[core] explicitly.
This is a comment-only change. Every changed line is a comment, and all
twenty-five GNU variants still assemble cleanly for their target core.
The fourteen files under ports/cortex_a{5,7,8,9,12,15,17} were regenerated
from ports_arch/ARMv7-A/threadx/common/src/tx_thread_system_return.S with
ports_arch/ARMv7-A/update.sh. The ARMv7-A SMP ports have no generator, so
those files were edited directly.
Fixes #734
Assisted-by: Copilot (Opus 5) <noreply@github.com>
|
||
|
|
c0aa4dbe29 |
Initialized the suspend status before the non-interruptable suspend in tx_thread_sleep, so a sleep no longer returns a stale error left over from an earlier timed-out suspension (#725)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
r52_fvp / r52 (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
When TX_NOT_INTERRUPTABLE is defined, _tx_thread_sleep did not set tx_thread_suspend_status to TX_SUCCESS before calling _tx_thread_system_ni_suspend. The interruptable path always did. Nothing else writes that field on behalf of a sleeping thread: _tx_thread_timeout resumes a TX_SLEEP thread directly and there is no suspend cleanup routine for sleep. The value therefore survived from whatever suspension the thread performed last, and tx_thread_sleep returned it. A tx_semaphore_get that timed out with TX_NO_INSTANCE would be followed by a tx_thread_sleep that also returned TX_NO_INSTANCE after sleeping correctly. The same change is applied to the SMP copy, which is identical. Assisted-by: Copilot (Opus 5) <noreply@github.com> |
||
|
|
70741198e9 |
Refused thread delete and reset while an exit transition is in progress (#724)
_tx_thread_shell_entry and _tx_thread_terminate both publish a thread's
terminal state -- TX_COMPLETED or TX_TERMINATED -- and then call that
thread's exit notification callback, before the thread has been detached
from the ready list and before either service has finished with the pointer
it holds to the control block. That terminal state is exactly the state
_tx_thread_delete and _tx_thread_reset accept as authorization to
invalidate or rebuild the control block, and neither service tested whether
the transition producing it had finished.
A callback could therefore delete the thread it was called for -- and then
lawfully recreate it over the same memory, since delete exists to permit
that -- while the scheduler was still linked to the old incarnation.
tx_thread_create zeroes the whole control block and can auto-start the new
one, so the old priority list is left heading at a block whose own priority
field names a different list, with the old priority's map bit set behind
nothing. Alternatively a callback could reset a terminated thread, which
moves it out of the terminal state, and then resume it: the interrupted-
suspension logic in _tx_thread_system_resume refuses to void a suspension
only while the state is still terminal, so with the reset allowed first the
resume clears the suspending flag and restores TX_READY, the outer service
then finds the flag clear and skips the removal, and tx_thread_terminate
returns TX_SUCCESS for a thread that is runnable again.
Interrupt masking does not close the window, because the kernel restores
the prior posture before invoking the callback deliberately. On SMP
TX_RESTORE also releases the global protection, so the target can be
executing on another core while its callback runs -- and a reset there
memsets the stack a live core is running on. On the Linux, Win32 and Win64
host simulation ports the consequence is more immediate than corruption:
TX_THREAD_DELETE_PORT_COMPLETION cancels and joins the host thread backing
the deleted thread, so a callback-side delete of the completing thread
destroys the host thread the callback is running on.
The fix marks the transition and has the two services refuse a marked
target, returning the errors they already document, TX_DELETE_ERROR and
TX_NOT_DONE. The refusal is transient and the same call succeeds once the
transition has completed, so no documented lifecycle is lost; and it is in
the core services rather than the _txe_ wrappers, so disabling error
checking cannot disable it. Refusing the reset is also what closes the
resume path, without touching _tx_thread_system_resume: its existing
terminal-state test is sufficient once nothing can turn the terminal state
into TX_SUSPENDED from inside the window.
tx_thread_suspending is the marker, rather than a new control-block field.
It already means "a suspension is in progress" and is already true across
the callback in the two interruptable paths, so no field is added, the
public structure is unchanged, and sizeof(TX_THREAD) is unchanged --
which matters, because the Module Manager's object handling depends on the
sizes of the control blocks. Widening its lifetime was checked against
every reader rather than assumed. There are four: two in
_tx_thread_system_suspend and two in _tx_thread_system_resume. In every
window this change widens, the state is TX_COMPLETED or TX_TERMINATED, and
both resume readers already refuse to void a suspension for exactly those
two states, so their behaviour is unchanged; and no suspension routine is
called on the target in those windows, so the suspend readers never see
them. No suspension-initiating service can set the marker again inside a
window either: every one of them acts on a thread that is ready or
suspended.
Three sites needed changing beyond the two refusals, and the shape of each
was decided by where the marker can safely be cleared:
- The non-ready branch of _tx_thread_terminate cleared the marker before
the terminated extension and the callback, which is what left them free
to act on a control block the service still had mutex-release
processing to do against. The clear moves to the common tail, after the
last dereference of the target, and becomes the single clear site for
the whole service. In the interruptable ready branch the flag is
already false there, because _tx_thread_system_suspend cleared it when
it detached the thread, so the tail store is a second store of a value
the flag already holds -- cheaper than testing for it, and it keeps one
clear site.
- Under TX_NOT_INTERRUPTABLE neither path set the marker at all, because
that configuration does not use the interruptable suspension path that
sets it. Both now set it before the callback. Interrupts being disabled
there does not help: the callback is reached by a direct call.
- In the TX_NOT_INTERRUPTABLE completion path the marker is cleared
before _tx_thread_system_ni_suspend rather than after it. That call
returns to the scheduler for a thread that is the current thread, which
a completing thread is, and does not come back; clearing afterwards
would leave a normally completed thread marked for ever and therefore
permanently undeletable. Nothing is lost by clearing early there,
because everything from that point to the detachment runs with
interrupts disabled and calls no application code.
The change is the same change twice. All four files are byte-for-byte
identical between common and common_smp at this commit and stay so after
it, so common_smp was written by copying rather than by repeating the
edits. tx_thread_system_suspend.c and tx_thread_system_resume.c, which do
differ between the kernels, are deliberately untouched.
Tests. The in-tree regression test goes to both trees and is byte-for-byte
identical between them. It drives seven scenarios: terminating a ready
non-current target with two peers ready at the same priority, with the
callback attempting the delete and recreating the block if it succeeded;
the same with the callback attempting the reset and then the resume;
terminating a target suspended on a semaphore while owning a mutex, which
is the non-ready branch; natural completion alone at its priority,
including the safe post-completion reset, terminate, delete and recreate at
another priority; self termination; a benign callback, whose notification
count and ordering are unchanged; and the state and boundary cases, where
the new refusal must not fire.
It measures rather than describes. The callback records the published
state, the marker, and the status of every lifecycle service it can reach,
calling the core service as well as the wrapper wherever a refusal is
expected. A snapshot taken under interrupt lockout -- which is the global
SMP protection on an SMP port -- checks that every ready list agrees with
the control blocks it heads, that the priority map agrees with the lists,
and that each execute pointer is a member of the list its own priority
field names. Every walk is bounded, so a corrupted ring costs an assertion
and not a hang, and no test in the suite can hang. Expectations are counted
inside a scenario and gated between scenarios, so a failing kernel reports
how much it failed by without being driven further into its own
corruption.
The consequences are demonstrated from the terminator's context rather than
the completing thread's, which is what makes the pre-fix behaviour an
assertion instead of a wedged simulator. Compiled against the unfixed
sources the test fails 9 of the 24 expectations it reaches in the
uniprocessor tree and 10 of 24 in the SMP tree, and the failures are the
finding: the callback-side delete succeeds, the recreate succeeds, the
consistency snapshot disagrees, the target is neither terminal nor detached
when the service returns, and a peer has left the ready ring the recreated
block hijacked.
The SMP tree gets a second test for the case that needs concurrency. The
victim is excluded to core 1 and spins there without relinquishing while
the controller, excluded to core 0, terminates it, so the callback runs on
one core while the target executes on another. The callback-side reset and
delete must both be refused, and a sentinel written into the unused low end
of the victim's stack must survive -- a reset would have memset the whole
stack before rebuilding the frame. Every wait is bounded, and if the remote
precondition cannot be established the test says so and drops only the
assertions that depend on it rather than reporting a pass it did not earn;
measured over twenty consecutive runs it established the precondition every
time.
TX_NOT_INTERRUPTABLE and TX_DISABLE_ERROR_CHECKING are not among the five
build configurations either tree compiles, and each tree builds the whole
library once per configuration, so neither can be reached from inside the
suites. The lines this change adds under TX_NOT_INTERRUPTABLE are therefore
in no configuration the trees build, and they are where the permanent-
undeletability failure mode lives, so they get their own harness rather
than a compile check: the four sources plus the two error wrappers are
compiled directly into a test executable, once per combination, with
recorders standing behind the scheduler services they call. That is what
makes the marker's value at the moment of detachment directly observable.
It runs 66 expectations under TX_NOT_INTERRUPTABLE, 66 under that with
error checking disabled, 45 under that with notification disabled, and 63
under error checking disabled alone; against the unfixed sources those fail
25, 22, 6 and 20 respectively. The harness lives in the uniprocessor tree
only, because the four sources are identical between the kernels and the
shim replaces the very primitive the SMP port differs in, so a second copy
would compile the same text under the same macros. It is deliberately left
out of the coverage instrumentation, since the same source under different
feature macros has a different line set and merging those would confuse the
union rather than add to it.
Results. Both suites pass in all five configurations with GCC 14: 103 of
103 in the uniprocessor tree, up from 98, and 116 of 116 in the SMP tree,
up from 114. Merged line coverage is 100% in the uniprocessor tree and
5172 of 5183 in the SMP tree, whose eleven uncovered lines are the same
eleven that were uncovered before this change and are in tx_byte_pool_search
and tx_thread_smp_utilities; all four changed files are at 100% line
coverage in both trees, and SMP branch coverage rises from 2819 of 3548 to
2831 of 3556. Cross-compiled with arm-none-eabi-gcc at -Wall -Wextra for
Cortex-M4 against common and for Cortex-A7 SMP against common_smp, all four
files produce no diagnostics at all and an identical warning set to before
the change, under -std=gnu99 and -std=c99 alike -- unlike the module ports,
-std=c99 does not fail on these base ports, and even -Wconversion is clean.
No MISRA deviation is required: explicit comparisons to TX_TRUE, existing
ThreadX types, single-entry and single-exit control flow, no goto, and two
added constant-time tests that change no real-time complexity.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
9ee198a2d6 |
Made the stack analyze binary search require several consecutive fill words, so unwritten holes in a used stack no longer cause the highest stack pointer to be under-reported (#720)
The binary search in _tx_thread_stack_analyze() accepted a probe location as unused as soon as a single word still held TX_STACK_FILL. A word inside an otherwise used region that simply was never written - the padding of a partially initialized local array, for example - therefore made the search move away from the real boundary and report far less stack usage than the thread had actually consumed, which in turn kept the stack guard from firing. The probe now walks down from the candidate location and requires TX_THREAD_STACK_ANALYZE_FILL_WORDS consecutive fill words before it treats the location as unused, stopping early at the lowest location already known to hold the fill pattern. The new macro defaults to eight words and can be overridden in tx_port.h; setting it to one restores the previous behavior. Holes shorter than the configured run no longer mislead the search, and the result is unchanged for stacks that contain no such holes. Assisted-by: Copilot (Opus 5) <noreply@github.com> |
||
|
|
d56f96f37f |
Refreshed the reason gcc_check leaves RISC-V out (#722)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
r52_fvp / r52 (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The header explained the exclusion by saying RISC-V "is not regressing" and that adding it would widen the toolchain download. The second half is still true; the first read as though nothing in CI exercised the family at all, which stopped being the case with #717. RISC-V is now the best-covered of the four excluded families rather than the least: regression_test.yml builds both ports and runs 955 tests on them under QEMU -- 475 on RV32 and 480 on RV64, across five build configurations each -- which is more than a compile-and-link check could establish. That is a stronger argument for leaving it out of this workflow than the original, so the sentence now makes it. The download figure is kept and quantified: the two bare-metal toolchains this workflow would have to fetch are about 500 MB apiece. Comment only; no behaviour change. scripts/check_gcc.sh names the same four families but states the exclusion without giving a reason for it, so it needs no matching edit. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
6c84e61d19 |
Gave install_riscv.sh the network hardening install.sh already had, and shared it between them (#721)
install.sh grew a retry loop, per-command timeouts and a deliberately non-gating apt-get update after this runner pool cost several whole runs: a mirror going silent for two hours, and a Hash Sum mismatch from a third-party repository the project does not even use turning builds red. Those lessons were local to that one file. install_riscv.sh had none of them, and #717 puts it on every pull request's critical path. Under set -e its bare apt-get update was a single point of failure for the whole suite -- the precise case install.sh downgrades to a warning on purpose -- and its two wget calls, each fetching about 500 MB, had no retry and no timeout. Rather than copy the helpers and let them drift again, they move to tx_ci_common.sh and both scripts source it, following the arrangement scripts/tx_windows_common.ps1 already uses on the Windows side. install.sh keeps its behaviour exactly: same APT_OPTIONS, same 120-second TIMEOUT, same three-attempt retry, and the comments explaining each of them travel with the code they explain. Two things are new: - TIMEOUT_LONG, 180 seconds, for a single large download. Sized against the 39 seconds each tarball took on 10 Sep 2026 and deliberately not larger: the install step is capped at ten minutes, and a per-attempt timeout able to swallow that cap would leave the retry loop no turn to take, which is the failure mode the apt comment already records. - fetch(), which verifies a SHA-256 before anything is unpacked. Both digests were taken from the releases API and then checked against the bytes the CDN actually serves. This is not an independent trust root -- expected value and file come from the same host -- but it pins the bytes, so a deleted and re-pushed tag or a replaced asset stops the build instead of being picked up silently. Also verifies qemu-system-riscv32 alongside riscv64. run.sh selects one per architecture, so both are worth failing on here rather than at the first test. Verified locally: retry returns 0 on success and 1 after three attempts; fetch accepts a correct digest and, on a wrong one, fails and removes the partial file; the source line resolves from the repository root, from an absolute path and through a symlink; and both recorded digests match the bytes served for the pinned tag. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
18abe103af |
Enabled the RISC-V regression suite in CI for QEMU targets (#717)
* Enable the RISC-V regression suite in CI The riscv: job in regression_test.yml has been present but commented out since the RISC-V CI infrastructure landed (scripts/install_riscv.sh, scripts/build_tx_riscv.sh, scripts/test_tx_riscv.sh, and the CMake tree under test/tx/cmake/riscv/). It has been gated on a note that read 're-enable when RISC-V CI is ready', with no other blocker recorded. The suite is ready. A local run of scripts/build_tx_riscv.sh followed by scripts/test_tx_riscv.sh against upstream/dev builds every RISC-V target and passes every registered test across all ten build configurations: RV32 (five configs): 5 * 95 = 475 tests, 0 failures, 0 timeouts. RV64 (five configs): 5 * 96 = 480 tests, 0 failures, 0 timeouts. The extra RV64 test is threadx_riscv_new_thread_fpu_state_test, which lives beside test/tx/cmake/riscv/regression/CMakeLists.txt and is added only when THREADX_ARCH is risc-v64 -- the RV32 stack builder still leaves the mstatus slot of the frame unwritten. Every test runs on qemu-system-riscv32 / qemu-system-riscv64 with -machine virt, and each configuration completes in roughly thirteen to fifteen seconds. The job is wired the same way the ThreadX, SMP and FreeRTOS suites are: it calls .github/workflows/regression_template.yml with the RISC-V install/build/test scripts and the RISC-V cmake_path, sets result_affix: RISC-V so its check name and artifacts are named, and carries skip_deploy: true because coverage publishing stays on the Linux suites for now. skip_coverage: true is kept because the CMake configurations under test/tx/cmake/riscv match the Linux suite's minus the coverage instrumentation, so gcovr has nothing to read -- the same reason the FreeRTOS lane sets it. The regression_test.yml triggers were not touched, so the RISC-V suite now runs on push and pull_request to master and dev alongside the other three suites. * Said why the RISC-V suite collects no coverage, instead of implying a missing build configuration The comment added with the job read "No coverage build configuration for the RISC-V suite yet". Both halves of what followed are true -- the configurations under test/tx/cmake/riscv are the Linux suite's minus the coverage one, and gcovr has nothing to read -- but "yet" points the next reader at a fix that would not work. Coverage here is not one missing default_build_coverage entry. These tests are bare-metal images run under QEMU with -bios none, and test/tx/cmake/riscv/bsp/syscalls.c has _write to the UART, _exit through the sifive_test device, and stubs for _close, _fstat, _isatty, _lseek, _read and _sbrk -- but no _open. gcov emits a .gcda by opening a path, so instrumenting these builds produces nothing regardless of how they are configured. Nor is the template's TX_COVERAGE=OFF what holds coverage off: test/tx/cmake/riscv is its own top-level project and never declares that option, so the value is inert there. Getting a figure out of this suite needs a transport off the target -- gcov's dump routines over the UART the BSP already drives, or semihosting. Worth stating plainly in a project that asks for 100% coverage, rather than leaving a reader to discover it after adding a configuration that cannot help. Comment only; the job's inputs are unchanged. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> --------- Co-authored-by: Frédéric Desbiens <frederic.desbiens@eclipse-foundation.org> |
||
|
|
42f1d2bcf0 |
Pointed the SMP coverage comment at the pull request that made its measurements (#719)
The paragraph explaining why the SMP floor sits at 99 rather than 100 ended with a bare reference that resolves to nothing a reader of this repository can open. It named a source outside the tree in place of one inside it. The real source is #677: it wrote the four tests that closed 53 of the 64 lines, took the measurements the paragraph quotes -- the 180,003 windows with zero handovers, and the three remaining lines in tx_thread_smp_utilities.c -- and raised the floor from 98 to 99. Citing it gives the next reader somewhere to go. This is the same correction #718 made across five other files; this line was missed because it sits in regression_test.yml rather than the template. Comment only; no behaviour change. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
39277cf026 |
Stated the compiler and coverage requirements directly, and said who the pinned toolchain default serves (#718)
* Stated the compiler and coverage requirements instead of citing a file no contributor can open
Seven comments across five files cited a maintainer-local document as the
source for two project requirements: that GCC 14 on Linux is the default
compiler, and that the coverage target is 100%. That document is not part of
this repository and is not published anywhere, so the citation gave a reader
nothing to follow -- it named a source they cannot open, in place of simply
stating the requirement.
Both requirements are real and both stay. Only the pointer goes: each comment
now states the requirement on its own terms, which is what the surrounding
prose was already doing everywhere else.
cmake/cortex_r52.cmake the pinned reference toolchain
scripts/check_gcc.sh why the script exists
.github/workflows/gcc_check.yml why the workflow exists, and the
GCC_VERSION pin
.github/workflows/r52_fvp.yml the GCC_VERSION pin
.github/workflows/regression_template.yml the coverage floor, twice
Comments only; no behaviour changes. Two paragraphs are rewrapped where the
shorter text left a ragged line. Verified that scripts/check_gcc.sh still
parses and prints its help from the header range it slices, and that
cmake/cortex_r52.cmake still configures the Cortex-R52 build.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Said who the pinned toolchain default in the CMake toolchain files serves
Both cmake/cortex_r52.cmake and cmake/cortex_m52.cmake default
ARM_TOOLCHAIN_PATH to a toolchains directory under the user's home, guarded by
an EXISTS check. Nothing said whether CI relies on that, and the natural
reading is that it does.
It does not. The three workflows that install a toolchain unpack it into the
workspace and cache it there, and r52_fvp.yml puts that directory on PATH
before configuring; scripts/check_gcc.sh passes -DARM_TOOLCHAIN_PATH at each of
its three CMake call sites. On a runner the guarded directory is absent, the
EXISTS check falls through, and the compiler comes from PATH. The default only
ever fires on a developer machine, where it is what makes a no-flag build work.
Both comments now say that, so the default is not mistaken for a CI dependency
and not removed as dead code. cortex_r52.cmake carries the explanation and
cortex_m52.cmake refers to it, matching the cross-reference already there.
The r52 comment also claimed absolute paths mean "the build does not depend on
PATH ordering", which is only true where the pinned directory exists -- in CI
the build depends on PATH and nothing else. Qualified accordingly.
Comments only; no behaviour changes. Verified that both toolchain files still
configure, and that the fall-through is real: with HOME pointed at a directory
holding no toolchains, cortex_r52.cmake configures against the arm-none-eabi-gcc
found on PATH.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
5e9d046b12 |
Compiled the module manager C sources with GCC, as clang already did (#716)
#689 added the module manager stage to check_clang.sh alone. The GCC half was never written, so the module manager C stayed unbuilt by the project's declared default compiler: 28 files of portable module manager under common_modules, plus the three to nine per-port files under ports_module/<core>/gnu/module_manager/src, across nine Arm module ports. The stage is deliberately check_clang.sh's, port for port and header for header, because a port covered by one check and not the other implies a parity the checks list does not have. The same two details the ports dictate carry over: an SMP port's control blocks come from common_smp rather than common, and the TrustZone ports need -mcmse for their cmse_nonsecure_entry functions to be honoured rather than ignored. The clang-only waiver does not: GCC implements the optimize attribute that tx_thread_secure_stack.c carries, so nothing needs suppressing for that file. One divergence is forced by the toolchain. txm_module_manager_absolute_load.c carries a #pragma message steering callers to the extended entry point, and the C stages treat any compiler output as a failure. check_clang.sh silences it with -Wno-#pragma-messages; GCC has no equivalent, and neither -Wno-pragmas nor any other -W option suppresses the note -- verified with 14.3.rel1. The note is therefore filtered out of the stage's output instead, together with the source quote GCC prints beneath it. The filter stops at the next line that begins a diagnostic of its own, so an error immediately following a waived note is still reported; that case is what the injected-defect run below checks. The workflow needed no trigger change: #689 added common_modules/** to both path lists in advance, for the stage that had yet to arrive. Its header comment is brought in line with what the workflow now runs. Verified with the toolchain CI pins, arm-gnu-toolchain 14.3.rel1, both triples. The full script passes and every port compiles every file, matching the counts check_clang.sh reports for the same nine ports: cortex_a35 31/31 cortex_a35_smp 31/31 cortex_a7 34/34 cortex_m0+ 33/33 cortex_m23 37/37 cortex_m3 33/33 cortex_m33 37/37 cortex_m4 33/33 cortex_m7 33/33 Verified that the stage fails as intended by injecting defects into a throwaway worktree: one in common_modules on the line straight after the waived pragma note, reported under all nine ports, and one in a single port's own source, reported only under that port. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
7f6e496233 |
Added the missing VFP enable field to the Cortex-R5/AC5 thread control block, which VFP builds were writing over the FileX pointer (#715)
The Cortex-R5/AC5 port defines TX_THREAD_EXTENSION_2 as empty while its assembly reads and writes the per-thread VFP enable flag at [thread, #144]. With TX_ENABLE_VFP_SUPPORT that offset lands on tx_thread_filex_ptr, so tx_thread_vfp_enable corrupted the FileX pointer and lazy save/restore tested an unrelated value. Defined TX_THREAD_EXTENSION_2 as ULONG tx_thread_vfp_enable, matching the Armv7-A ports and the Cortex-R4/R5 GNU and AC6 ports fixed earlier, and added tx_port_offset_check.c so a future layout change breaks the build instead of silently retargeting the accesses. The offset was measured at 144 for this port and asserted. Fixes #382 Assisted-by: Copilot (Opus 5) <noreply@github.com> |
||
|
|
a8c84593dd |
Added ARMv7-A SMP Linux build scripts(A5,A7,A9) (#674)
* ports: add Linux build scripts for Cortex-A5 SMP * ports: fix Cortex-A5 SMP sample linking on Linux * ports: add Linux build scripts for Cortex-A7 and A9 SMP * ports: fix Cortex-A7 SMP assembly source path * wip: add ATFE selection to Armv7 SMP scripts * ports: complete Armv7 SMP Linux toolchain support * ports: add local startup for Armv7 SMP samples * ports: add license headers to Armv7 SMP scripts * Dropped the preprocessing flag the file extension now carries Two lines named assembly sources that #672 renamed. That change moved twenty-nine files under gnu trees from .s to .S, because GAS runs the C preprocessor on .S and not on .s: in a .s file every # line is a comment, so a #define is never substituted and an #if/#else pair emits both arms. Four files were silently doing the wrong thing as a result, including ports_smp/cortex_a7_smp/gnu/src/tx_thread_smp_unprotect.s, which ignored all four of its own feature macros. These scripts had worked around the same defect with -x assembler-with-cpp rather than hitting it, which was correct when they were written. With the rename the flag is redundant and the lowercase names no longer resolve, so cortex_a5_smp/build_threadx.sh failed with "cc1: fatal error: tx_initialize_low_level.s: No such file or directory". The other seven references to those two files across these three scripts already named them with a capital S. Verified with the pinned Arm GNU 14.3.rel1 rather than the 13.2 that a distro package supplies: build_threadx.sh and build_threadx_sample.sh both succeed for a5, a7 and a9, all three link a sample_threadx.out, and scripts/check_gcc.sh passes end to end with its example stage reading 45 of 45, up from 42. Assisted-by: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Frédéric Desbiens <frederic.desbiens@eclipse-foundation.org> |
||
|
|
383dd311cd |
Removed the vector table offset register and system stack pointer setup from the Cortex-M low-level initialization (#714)
_tx_initialize_low_level wrote VTOR and derived _tx_thread_system_stack_ptr from the reset vector on every Cortex-M port. Programming VTOR is the job of the low-level startup code that runs before the kernel is entered, and doing it again inside ThreadX silently overrode a vector table that the application, a bootloader or a firmware update had already installed. It also forced every application to export a vector table symbol that the kernel itself never needed. _tx_thread_system_stack_ptr is never read by any Cortex-M port, so the value copied out of the reset vector served no purpose. Both blocks are removed from all Cortex-M ports, together with the now-unused symbol declarations. The GreenHills files keep the NVIC base address load that the removed block used to leave in r0 for the SysTick setup that follows. Applications that relied on ThreadX programming VTOR must now set it in their startup code. The bundled examples link their vector table at address zero and run on the reset default. Fixes #370 Assisted-by: Copilot (Opus 5) <noreply@github.com> |
||
|
|
850a172bac |
Added lazy FPU stacking and QEMU functional tests for RV64 GNU port (#549)
* add lazy FPU stacking to context save/restore Save mstatus/sstatus to stack slot 29 and skip floating-point register save/restore when FS is Off (bits 14:13). This avoids unnecessary FP context work for threads that do not use the FPU. - context_save: check FS in nested and first-level interrupt paths - context_restore: gate FP restore on nested, no-preempt, and preempt paths - use sstatus when TX_RISCV_SMODE is defined, otherwise mstatus * add QEMU virt CMake build and automated test runner Wire the QEMU virt demo into the CMake build system and add a Python/GDB functional test runner, mirroring the risc-v32/gnu port. - Add qemu_virt/CMakeLists.txt to build kernel.elf and register the check-functional-riscv64 target (requires Python3; skipped if absent) - Link kernel.elf with --whole-archive so all ThreadX symbols resolve - Pin _start at 0x80000000 via .text.boot in entry.s and KEEP(*(.text.boot)) in link.lds - Extend demo_threadx.c with fpu_test_val and shorten thread_0 sleep for GDB-driven FPU, timer, and preemption checks - Add test/azrtos_test_tx_gnu_riscv64_qemu.py; verified passing on QEMU virt (FPU, timer interrupt, preemption) * Clean up RV64 PR scope and remove QEMU test integration leftovers Revert accidental RV64 qemu_virt test/CMake integration changes and keep this branch focused on lazy FPU context handling only. Also remove unintended TX_RISCV_SMODE-based mstatus/sstatus save path and align comments/logic to mstatus-only behavior. * Initialize mstatus.FS in RV64 stack build so new threads start with clean FP state Slot 29 was left uninitialized while context restore reads it as an FP-live hint; garbage FS bits could make a new thread inherit the previous thread's floating-point registers. * Add RV64 regression test for the FP state of a newly created thread The test dirties every floating point register, then creates a thread and checks that the stack builder wrote the mstatus slot and that the new thread starts with all floating point registers zeroed. It is registered for RV64 only, since the RV32 stack builder still leaves the slot unwritten. * Completed the RISC-V64 lazy FPU so the restore side matches the save side The lazy FPU save in this branch skips the floating-point stores when mstatus.FS is Off, and records the mstatus it judged that on in frame slot 29. Merged onto current dev, only the save side had that treatment: both restore paths and the scheduler's interrupt-frame path still reloaded the FP registers unconditionally, from slots the save had deliberately left alone. A thread that never touched the FP unit would have had whatever the frame happened to contain loaded into its registers, and FS driven to Dirty on the way out. The guard is added at the three places that consume an interrupt frame: both paths in _tx_thread_context_restore, and _tx_thread_schedule_loop. Each reads slot 29 and skips the FP block when FS was Off, which is the same shape the risc-v32 port already uses. The solicited path is deliberately left alone. _tx_thread_system_return saves the callee-saved FP registers unconditionally, so restoring them unconditionally is consistent; making that pair lazy as well is a separate change, and risc-v32 is the model for it. Verified with QEMU on all five configurations: risc-v64 96 of 96 passing, five configurations, nothing unlinkable risc-v32 95 of 95 passing, five configurations, unchanged functional check-functional-riscv64 passes every check The ninety-sixth test is the one this branch adds. It is load bearing: seeding stack build with FS = Off instead of Initial makes it fail, and restoring the seed makes it pass, so it guards the behaviour the rest of this branch is about rather than passing regardless. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> --------- Co-authored-by: r <r@r> Co-authored-by: Frédéric Desbiens <frederic.desbiens@eclipse-foundation.org> |
||
|
|
16297d7a2a |
Added a workflow that runs the Cortex-R52 images on the Armv8-R AEM FVP (#688)
* Added a workflow that runs the Cortex-R52 images on the Armv8-R AEM FVP Nothing in CI executed a single instruction of any ThreadX port. gcc_check compiles and links -- its own header says it "executes nothing" -- and its CMake stage covers one R52 configuration, the base FVP example. A change that assembles cleanly, links cleanly and then hangs on the first context switch passed every required check. This workflow builds both supported configurations and runs them on the Armv8-R AEM FVP, asserting each image's self-reported result. The feature configuration matters as much as the default one: VFP with the hard float ABI, FIQ, IRQ nesting and FIQ nesting gate whole assembly blocks that the default configuration never assembles, and it builds eight images where the default builds five. Arm distributes the model free of charge but behind a click-through licence with no stable unauthenticated URL -- the developer.arm.com permalink forms all 404 for it -- so the download location is a repository variable rather than a literal. When it is unset the build lanes still run and still gate the pull request; only execution is skipped, and it says so in the log and in the job summary rather than passing quietly. The image list is read from the generated ninja graph rather than kept in the workflow. The images are EXCLUDE_FROM_ALL, so a bare build reports "no work to do", and reading the graph means a target added to CMakeLists.txt cannot escape the check by nobody remembering to list it here. The module manager port and the S32Z280 targets are not on dev yet. Their lanes belong in this workflow when they land, not in a second one. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> * Kept the FVP cache honest and stopped ctest passing on an empty suite Two corrections to the new workflow. The FVP cache key named the URL and hashed the workflow file. hashFiles() reads files, so it cannot hash a repository variable, and pointing it at this workflow keyed the cache on the file's own contents instead. That is wrong in both directions: repointing FVP_AEMV8R_URL at a different model kept serving the old one from cache, which is exactly what the comment claimed the key prevented, and editing anything in this file forced a needless refetch of a large archive. A small step hashes the URL and the key uses that, so it now tracks what it names. ctest reported success when it found nothing to run. Its default is to pass an empty suite, and the example registers its tests inside both if(FVP found) and if(Python3_FOUND). Either one going unmet on a runner would have left this stage green having executed no image at all, which is the hole the workflow exists to close. --no-tests=error makes an empty suite a failure. Verified with the pinned toolchain, Arm GNU 14.3.rel1: both configurations configure, enumerate and build, and the image lists are the ones claimed. default: 5 images boot_check demo_m2 demo_m3 demo_mpu demo_threadx feature: 8 images the above plus demo_fiq demo_m5 demo_nesting Also checked that the target enumeration survives the real ninja graph: the generated graph offers twenty .elf paths, and the filter reduces them to those eight, dropping the cmake_object_order_depends_target aliases and the CMakeFiles copies. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> --------- Co-authored-by: r <r@r> |
||
|
|
8be4c45cce |
Compiled the module manager C sources, which no check had ever built (#689)
* Compiled the module manager C sources, which no check had ever built #672 corrected this script's assembly glob and brought the module ports into the count, but only their assembly. Their C stayed outside every check: 27 files of portable module manager under common_modules, plus the per-port code under ports_module/<core>/gnu/module_manager/src. By this script's own standard -- a port simply absent from the count reads as covered -- 293 files across nine Arm module ports were compiled by nothing, with either compiler. Each module port ships its own tx_port.h and txm_module_port.h carrying the control-block extensions the dispatch code needs, so a port is compiled against its own headers rather than the base port's. Two details the ports themselves dictate: An SMP port's control blocks come from common_smp. Pairing cortex_a35_smp with the single-core headers hid _tx_thread_smp_protect and _tx_thread_smp_unprotect behind implicit declarations and lost tx_thread_smp_core_executing from TX_THREAD, so fourteen files reported errors for a port that builds correctly. The TrustZone ports carry cmse_nonsecure_entry, which needs -mcmse to be honoured rather than ignored. tx_thread_secure_stack.c also carries GCC's optimize attribute, which clang does not implement; that divergence is suppressed by name, for a file GCC builds cleanly. All nine ports compile: 293 of 293. Verified that the stage fails as intended by injecting a defect into a throwaway worktree -- a defect in common_modules is reported under every port, one in a port's own source only under that port. Assisted-by: Claude Code (Opus 5) * Ran the module manager stage against the sources that reach it Three corrections to the new stage, all found by running it against dev rather than against the tree it was written on. The workflows did not trigger on common_modules. Both check_clang.yml and check_gcc.yml list ports_module but not common_modules, and the portable module manager under it is the larger half of what the stage compiles -- 28 of the 31 to 37 files each port builds, plus two include directories. A change there left the stage unrun, which is the same "absent from the count reads as covered" that the stage exists to close. Both lists gain common_modules, and they stay identical to each other as the comment in each asks. A deliberate deprecation notice read as a build failure. Since this PR was opened, txm_module_manager_absolute_load.c gained a #pragma message steering callers to the extended entry point. The stage treats any compiler output as a failure, so that one notice failed every port: nine failures on a tree where nothing is wrong. Pragma messages are now waived for the stage, because a notice to callers is not a defect in the file that carries it. That failure also printed nothing. Both C stages report by grepping the output for "error:", so a diagnostic that is not an error produced a bare FAIL line with no reason under it, and the only way to learn the reason was to reproduce the compile by hand. Both stages now fall back to showing what the compiler actually said. The TrustZone attribute waiver is narrowed to the one file that needs it. tx_thread_secure_stack.c carries GCC's optimize attribute, which clang does not implement; it is the only file among the 300-odd this stage compiles that does. Waiving the warning for the whole port would have swallowed a stray unknown attribute anywhere else in it. Verified with the same toolchain CI uses, ATfE 22.1.0: the full script passes, and every port compiles every file. cortex_a35 31/31 cortex_a35_smp 31/31 cortex_a7 34/34 cortex_m0+ 33/33 cortex_m23 37/37 cortex_m3 33/33 cortex_m33 37/37 cortex_m4 33/33 cortex_m7 33/33 The counts are each one higher than this PR first reported, because txm_module_manager_absolute_load_extended.c has landed since. The stage picked it up with no edit, which is what globbing the directories was for. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> --------- Co-authored-by: r <r@r> |
||
|
|
6ce8d5cc76 |
Marked every published ThreadX include directory as SYSTEM so applications no longer get warnings from ThreadX headers (#713)
* Marked every published ThreadX include directory as SYSTEM so applications no longer get warnings from ThreadX headers
Commit
|
||
|
|
4d90a0c21c |
Stopped the win32 and win64 ports from enabling performance metrics and event trace (#676)
* win32: do not always enable trace or performance metrics in tx_port.h if these are required they can be enabled in tx_user.h * win64: do not always enable performance metrics in tx_port.h if required they can be enabled in tx_user.h * Removed the disabled blocks rather than commenting them out The win32 and win64 ports were the only two that turned performance metrics on for the application, and win32 the only one that turned event trace on. Leaving those to tx_user.h is right: the symbols extend the control blocks, so a port that sets them behind the application's back changes structures the application also sees. The blocks were disabled with #if 0 rather than deleted. That is the form MISRA C:2012 Directive 4.4 is about -- sections of code should not be commented out -- and it leaves two copies of a list that now has no reader. They are removed, and a short note in their place says where the symbols belong and why the port does not set them. No behaviour change beyond what this pull request already made. Checked that the preprocessor nesting in both headers is still balanced. Worth recording for whoever looks next: with this in, no port defines either symbol. The linux port carries the same list commented out, which reads at a glance like a third case but is not one. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> --------- Co-authored-by: r <r@r> Co-authored-by: Frédéric Desbiens <frederic.desbiens@eclipse-foundation.org> |
||
|
|
40db27e843 |
riscv32: spec compliance and regression test fix (#691)
* riscv32: spec compliance and regression test fix Signed-off-by: Akif Ejaz <akifejaz40@gmail.com> * Derived the RISC-V32 frame sizes from the port contract in one place tx_port.h published TX_RISCV_TRAP_FRAME_SIZE for the GNU BSP assembly, but nothing in the port consumed it. Six .S files each rebuilt the same numbers from their own #if, so the interrupt frame size was written out in seven places and the solicited frame size in three. That is the shape that produced the RISC-V64 fault fixed in #708, where the port moved to a padded frame and one copy of the constant did not. The sources now include tx_port.h and take both sizes from it, and no literal frame size remains in the port. TX_RISCV_SOL_FRAME_SIZE joins the contract, since the solicited frame was never published at all. The emitted code is unchanged: 400 and 176 bytes for ILP32D, 128 for soft-float, confirmed by disassembly before and after. Two further corrections: _tx_initialize_low_level carried .global immediately followed by .weak, so the symbol stayed weak and the .global did nothing. Weak is what the port wants, because the example and regression BSPs both provide their own definition, so the stray .global is removed rather than the .weak. Verified with nm that the symbol is still W. The QEMU runner seeded fpu_verified from skip_fpu, so a soft-float run satisfied the FPU gate whether or not the script ever reported the skip. It now starts false and is set only when the skip marker is present, so a run that dies before reaching that point fails instead of passing. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> --------- Signed-off-by: Akif Ejaz <akifejaz40@gmail.com> Co-authored-by: Frédéric Desbiens <frederic.desbiens@eclipse-foundation.org> |
||
|
|
3fc28d8979 |
Allowed a BASEPRI-masked interrupt to wake the Cortex-M idle loop from WFI (#711)
The idle loop in _tx_thread_schedule masks interrupts before re-reading _tx_thread_execute_ptr, so that an interrupt cannot make a thread ready between that read and the WFI and then be lost. With PRIMASK that is safe: the architecture excludes PRIMASK from the WFI wake-up condition, so the interrupt still wakes the processor and is taken as soon as the loop re-enables interrupts. BASEPRI is not excluded. When TX_PORT_USE_BASEPRI is defined, the mask written at the top of the loop is still in place at the WFI, so any interrupt at or below TX_PORT_BASEPRI fails to wake the processor at all. On a system where every interrupt is managed by ThreadX, that is every interrupt, and the processor stays in WFI or in the low power mode entered by tx_low_power_enter until something outside the mask happens. Set PRIMASK and clear BASEPRI for the duration of the WFI, then restore BASEPRI and clear PRIMASK. Interrupts remain masked across the whole window, so the original race is still closed, but the wake-up condition is now evaluated with BASEPRI clear and any enabled interrupt can end the wait. A pending interrupt is taken at the CPSIE i, or at the existing unmask further down the loop if it falls below TX_PORT_BASEPRI. The change is made in the ARMv7-M and ARMv8-M sources under ports_arch, including the module manager, and propagated to the generated ports with the copy scripts. The ghs and keil ports do not implement TX_PORT_USE_BASEPRI, and Cortex-M0/M0+/M23 have no BASEPRI register, so those ports are unaffected. Verified by assembling every patched GNU port for Cortex-M3, M4, M7, M33, M55 and M85, with and without TX_PORT_USE_BASEPRI, TX_LOW_POWER and TX_ENABLE_EXECUTION_CHANGE_NOTIFY, and by checking the disassembly of the idle loop. Fixes #279 Assisted-by: Copilot (Opus 5) <noreply@github.com> |
||
|
|
146d57b235 |
Fixed the RISC-V64 trap frame size mismatch in the regression test BSP (#708)
The RISC-V64 port moved its interrupt frame to 528 bytes (65 slots plus 8
bytes of padding, so sp stays 16-byte aligned at a call) and published the
size as TX_RISCV_TRAP_FRAME_SIZE. The port sources were converted to use
it, but the shared regression test BSP was not: its trap_entry still
allocated a hardcoded 65 * REGBYTES, or 520 bytes.
Every interrupt therefore unwound 8 bytes more than it allocated:
trap_entry: addi sp,sp,-520
_tx_thread_context_restore: addi sp,sp,528
On the RISC-V64 regression suite that left 24 of 95 tests failing in the
default configuration, typically as an illegal instruction once execution
reached a corrupted frame. The example BSP under the port directory was
converted with the port and was unaffected, which is why the functional
QEMU test kept passing.
The test BSP now takes both frame sizes from the port it is linked
against, so the two cannot drift apart again. A port that publishes no
contract keeps the historical layout, so the RISC-V32 side is unchanged
until its own port publishes one.
TX_RISCV_TRAP_CALL_FRAME_SIZE is restored to the RISC-V64 tx_port.h. It
was removed as unused when the frame sizes were introduced, but it is
part of the same contract: it is the space a trap entry reserves around a
call into C, and the psABI requires 16 bytes there rather than one
register slot.
Verified on QEMU with every linkable test built, comparing against the
commit before the port change:
before the port change 2 failures out of 95 (both unlinkable)
current dev 24 failures out of 95
with this change 2 failures out of 95 (both unlinkable)
The two remaining failures predate all of this: newlib pulls _impure_ptr
out of R_RISCV_HI20 range for time(), so those two binaries do not link.
RISC-V32 is unchanged at 2 failures across all five configurations.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
9a03838381 |
Stopped the test runners from reporting an incomplete build as test failures (#709)
The cmake test runners drove Ninja with its default keep-going of 1, so the
first failing target ended the build. Every target scheduled after it was
simply absent, and ctest reports a missing binary as a failing test. A
single link error therefore produced a failure count that moved with build
scheduling order rather than with the code.
Measured on the RISC-V64 regression suite, where two targets genuinely
cannot link:
before 79 of 95 test binaries built
after 93 of 95 test binaries built
Fourteen perfectly good binaries were being skipped and counted as
failures. Passing -k 0 lets Ninja finish everything it can; the build still
exits non-zero when a target fails.
Three related problems in the same paths are fixed with it.
A failing configuration used to abort the loop over configurations, so
under set -e the ones after it went unbuilt or untested. The build loops
and the serial test loops now accumulate status and return it at the end,
which is what the parallel test branch already did with wait, and what
cmake_bootstrap.sh already documented for ctest.
Capturing that status removes the set -e protection inside the functions,
so two latent faults become reachable and are closed here. A failed pushd
would have let ctest run in the source tree, where it finds no tests and
reports success; the pushd is now guarded. And ctest's status was
discarded by the popd that follows it, so a configuration with failing
tests returned 0 and was reported as a pass; the status is now carried
past the popd and the summary steps.
Verified on the RISC-V32 suite, which has two genuine failures in each of
its five configurations. Both the serial and the parallel branch now test
all five and exit 8, where the serial branch previously stopped after the
first configuration.
The tx and smp runners are symlinks to scripts/cmake_bootstrap.sh, so they
are covered by the one change there.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
37f48fa83c |
Removed FIFO queueing from the ARMv7-A SMP ports so an ISR can no longer deadlock waiting for protection (#707)
The Cortex-A5, A7 and A9 SMP ports guarded the inter-core protection with a FIFO wait list: a core that could not take the lock added itself to a queue, incremented _tx_thread_smp_protect_wait_counts[core], and only the core at the head of the queue was allowed to acquire. Two parts of that scheme require the waiting core to be interruptible. _tx_thread_smp_unprotect() refuses to release the protection while the releasing core's own wait count is non-zero, and a queued core is taken back out of the list by _tx_thread_context_restore() when it is preempted. Neither can happen on a core that is spinning inside an ISR with interrupts already masked, because the spin loop restores the caller's interrupt posture rather than enabling interrupts. The core stays in the list forever and the system deadlocks with cores stuck in the wait loop. This is the deadlock Microsoft removed from the ARMv8-A SMP ports in 6.1.11, by dropping the wait list and using a plain LDAXR/STXR spinlock. The same removal was announced for the ARMv7-A ports at the time but was never made, so those three ports have carried the deadlock ever since. The FIFO queueing is now removed from the ARMv7-A SMP ports as well. _tx_thread_smp_protect() becomes an LDREX/STREX spinlock that releases interrupts between attempts, matching the sequence already used by the Cortex-R8 SMP port; _tx_thread_smp_unprotect() no longer consults the wait counts; and _tx_thread_context_restore() no longer has to unqueue a preempted core. The now unreferenced wait list macro headers are deleted. Verified by assembling all three GNU ports with arm-none-eabi-gcc for cortex-a5, cortex-a7 and cortex-a9, with and without TX_ENABLE_FIQ_SUPPORT, TX_ENABLE_WFE and TX_MPCORE_DEBUG_ENABLE, and by reading back the disassembly of the new protect sequence. The port consistency checks pass. Fixes #219 Assisted-by: Copilot (Opus 5) <noreply@github.com> |
||
|
|
d3fb72b6dd |
Aligned the simulator ports' fake stack pointer so ThreadX no longer performs misaligned ULONG accesses (#705)
The Linux, Win32 and Win64 simulation ports build a fake initial stack pointer by subtracting a fixed 8 bytes from tx_thread_stack_end. That field addresses the last byte of the thread's stack area, so it is one less than an aligned address and the resulting pointer is misaligned by construction, no matter how well aligned the stack the application supplied was. Two ULONG accesses then use that pointer. _tx_thread_stack_build() itself clears the word below it, and _tx_thread_create() copies it into tx_thread_stack_highest_ptr, which TX_THREAD_STACK_CHECK dereferences on every suspend and resume when TX_ENABLE_STACK_CHECKING is defined. Both are undefined behaviour. They happen to work on x86 but are reported by GCC's undefined behaviour sanitizer, and would fault on a host that requires natural alignment. The fake stack pointer is now rounded down to a ULONG boundary, which leaves it inside the stack area and makes both accesses aligned. Verified by building the Linux port with -fsanitize=undefined and running the demo: the two reported diagnostics are produced before the change and neither appears after it. The tx and smp regression suites pass, 98 and 114 tests respectively. Fixes #218 Assisted-by: Copilot (Opus 5) <noreply@github.com> |
||
|
|
2e48ce0d1b |
Removed the stray build artifacts committed with the RISC-V64 port fix (#706)
A local CMake build tree was staged by mistake alongside the RISC-V64 spec compliance work in #698, putting nine generated files on dev, including a compiled kernel.elf, build.ninja, the Ninja dependency logs and a QEMU run log. None of it belongs in the repository. The root .gitignore listed build directories by name rather than by pattern, so build/, build_qemu/, build_m7/ and the build_r52 variants were covered but a differently named tree was not. Those entries are replaced with a single build*/ pattern, which covers every existing name and any future one. No tracked file matches the new pattern. The artifacts remain reachable in history; only the working tree is corrected, since rewriting a shared branch is the greater harm. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
164f211a01 |
riscv64: spec compliance and regression test fix (#698)
* spec compliance Signed-off-by: Akif Ejaz <akifejaz40@gmail.com> * revert the demo changes Signed-off-by: Akif Ejaz <akifejaz40@gmail.com> * Restored the FPU demo hook so the RISC-V64 functional test can pass The new check-functional-riscv64 target verifies FPU context switching by watching fpu_test_val advance by 1.1f on each pass through thread_6_and_7_entry. The GDB script deliberately treats a missing symbol as a failure rather than silently skipping the check, but the demo no longer defined it, so the target failed on every run: FPU_VERIFIED_FAIL_NO_SYMBOL The definition and the increment are restored, matching what the risc-v32 demo already carries. The functional target now passes end to end. Three small corrections are folded in: - tx_port.h carried a comment stating that the ISA string must include Zicsr, but nothing enforced it, so an rv64imac build failed with a wall of assembler "unrecognized opcode" errors. It now stops at one clear diagnostic. - Removed TX_RISCV_TRAP_CALL_FRAME_SIZE, which nothing referenced. - The example .gitignore listed qemu-riscv32.log, but the runner writes qemu-riscv64.log, so the generated log showed up as an untracked file. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> --------- Signed-off-by: Akif Ejaz <akifejaz40@gmail.com> Co-authored-by: Frédéric Desbiens <frederic.desbiens@eclipse-foundation.org> |
||
|
|
515ab8aba1 |
Added the missing memory barrier so ARMv7-A SMP schedulers no longer miss preemptions (#704)
_tx_thread_schedule stores the newly selected thread into _tx_thread_current_ptr[core] and then reloads _tx_thread_execute_ptr[core] to detect a concurrent scheduling decision made by another core. On the other side, _tx_thread_smp_core_interrupt stores the execute pointer and then reads the current pointer to decide whether an inter-core interrupt is required. This is a store-buffer pattern: without a barrier both cores can observe stale values, the interrupt is skipped, and a ready thread with the highest priority is never scheduled. The barrier was added to the ARMv8-A SMP scheduler in 6.2.1, but the ARMv7-A SMP ports were left untouched even though they implement the same protocol. This adds the corresponding DMB to the Cortex-A5, Cortex-A7 and Cortex-A9 SMP schedulers for both the AC5 and GNU toolchains. The GNU variants were verified by assembling them with arm-none-eabi-gcc 13.2.1 for their respective cores. Refs #209 Assisted-by: Copilot (Opus 5) <noreply@github.com> |
||
|
|
e99f0d4207 |
Corrected the module kernel stack size so it no longer overstates the usable stack (#701)
The module manager recorded tx_thread_module_kernel_stack_size as the raw TXM_MODULE_KERNEL_STACK_SIZE constant, but the end of the kernel stack is aligned downwards to an eight-byte boundary while _txm_module_manager_object_allocate only guarantees ULONG alignment. The recorded size could therefore overstate the usable stack by up to seven bytes. The scheduler copies this value into tx_thread_stack_size whenever a user mode module thread enters the kernel, so the overstated value is visible to RTOS-aware debuggers and to anything built on it. The size is now derived from the aligned end minus the start. Also documented that TX_ENABLE_STACK_CHECKING is not supported for module threads. Refs #181 Assisted-by: Copilot (Opus 5) <noreply@github.com> |
||
|
|
945f5f5caa |
Moved the ARC ISR enter callout onto the system stack (#700)
_tx_thread_context_save() calls _tx_execution_isr_enter when TX_ENABLE_EXECUTION_CHANGE_NOTIFY is defined. On the path where an interrupt preempted a running thread, that call was made before the switch to the system stack, so the 32-byte call frame and the whole stack footprint of the callout were taken from the interrupted thread's stack, on top of the 160-byte interrupt frame the port had already allocated there. The callout is supplied by the application, so its stack usage is not bounded by ThreadX, and it is charged to every thread that happens to be running when an interrupt arrives. The two other callout sites in the same routine, the nested save and the idle system save, already run on the system stack, as does the _tx_execution_isr_exit call in _tx_thread_context_restore. _tx_thread_schedule was reordered in 6.1.9 so that _tx_execution_thread_enter runs on the system stack rather than the thread stack; the same reorder was never applied to the context save. The switch to the system stack now happens before the callout in the ARCv2_EM, ARC_HS and SMP ARC_HS ports. _tx_thread_context_fast_save is unchanged because the fast interrupt path never switches stacks by design. Fixes #149 Assisted-by: Copilot (Opus 5) <noreply@github.com> |
||
|
|
29afcc3946 |
Fixed a kernel stack leak when deleting user-mode module threads (#692)
* modules: free kernel stack on thread deletion Signed-off-by: Prashit Vora <prashitvora2006@gmail.com> * Preserved the thread object release when the kernel stack cannot be freed Releasing the kernel stack ahead of the thread object made a failure of the kernel stack deallocation abort the thread object release. The thread had already been deleted at that point, so the thread object would have stayed allocated for the lifetime of the module. The thread object is now always released once the delete succeeds, and the kernel stack failure is reported only when it does not mask a thread object failure. --------- Signed-off-by: Prashit Vora <prashitvora2006@gmail.com> Co-authored-by: Frédéric Desbiens <frederic.desbiens@eclipse-foundation.org> Assisted-by: Copilot (Opus 5) <noreply@github.com> |
||
|
|
b0ec8bfbb9 |
Fixed the invalid module data pointers in the absolute module load (#699)
An absolutely located module has its code and its data placed at two independent fixed addresses by the module's linker script. The module preamble carries the code and data sizes but not the data address, so _txm_module_manager_absolute_load() could not determine where the module's data area was. It computed txm_module_instance_data_start from the code size and the preamble size, which yields a size rather than an address, and it set txm_module_instance_module_data_base_address one past the end of the byte pool allocation. Added _txm_module_manager_absolute_load_extended(), which accepts the module's data area address from the caller. Deprecated _txm_module_manager_absolute_load(), which now forwards to the extended service with an unknown data area location and rejects modules that request memory protection, since the memory protection hardware cannot be programmed to cover an unknown data area. Fixes #450 Assisted-by: Copilot (Opus 5) <noreply@github.com> |
||
|
|
d7789f0b12 |
Fixed the clobbered return address in the RISC-V context save (#696)
_tx_thread_context_save() returns to its caller with ret, which uses the return address held in ra. When TX_ENABLE_EXECUTION_CHANGE_NOTIFY was defined, the call to _tx_execution_isr_enter overwrote ra with the address of the instruction following the call, so the subsequent ret returned into _tx_thread_context_save itself instead of the interrupt service routine. The return address is now saved on the stack around the call and recovered afterwards, which is the same idiom already used by the Arm ports. The fix covers all three affected paths (nested save, thread save and idle system save) in the risc-v32 GNU, risc-v32 IAR and risc-v64 GNU ports. Fixes #348 Assisted-by: Copilot (Opus 5) <noreply@github.com> |
||
|
|
508af549da |
Fixed the incorrect loop bound constant in the IAR file lock support (#695)
The IAR multithreaded library support code allocates its file lock mutexes from an array of _MAX_FLOCK entries, but the wrap-around check and the exhaustion check in __iar_file_Mtxinit() both compared against _MAX_LOCK, the bound of the unrelated system lock mutex array. When _MAX_FLOCK is greater than _MAX_LOCK, the free mutex index wrapped early and the exhaustion check reported failure while free entries remained, so *m was set to TX_NULL and the application faulted the first time a file lock was taken. When _MAX_FLOCK is smaller than _MAX_LOCK, the free mutex index was allowed to run past the end of __tx_iar_file_lock_mutexes and the exhaustion check could never fire. Corrected all four comparisons in each of the 27 copies of tx_iar.c. Fixes #444 Assisted-by: Copilot (Opus 5) <noreply@github.com> |
||
|
|
17ff21a07f |
Fixed the garbled comment in the POSIX condition variable sources (#694)
The comment above the internal semaphore lookup in the POSIX condition
variable implementation was garbled: it contained a capitalisation typo
("COndition") and an incomplete sentence ("into a semaphore a cast").
The four affected sources now share a single, correct wording.
This is a comment-only change; there is no functional impact.
Fixes #419
Assisted-by: Copilot (Opus 5) <noreply@github.com>
|
||
|
|
c469a7756b |
Fixed the missing immediate prefix on MOV in Cortex-M schedulers (#693)
The BASEPRI-masking path in tx_thread_schedule wrote "MOV r0, 0" rather than "MOV r0, #0". UAL requires the "#" prefix on an immediate operand. GNU as and the LLVM-based assemblers accept the unprefixed form and emit the intended encoding, but stricter assemblers reject it outright, so the affected ports could not be built with those toolchains. The GNU and AC6 sources had already been corrected; this brings the IAR and AC5 sources into line. Verified that both spellings assemble to the same Thumb-2 encoding (f04f 0000), so this is a source-correctness fix with no change in generated code or runtime behaviour. Covers 31 occurrences across the Cortex-M3, M4, M7, M33, M52, M55 and M85 ports, their module manager counterparts, and the shared ARMv7-M and ARMv8-M architecture sources. Fixes #461 Assisted-by: Copilot (Opus 5) <noreply@github.com> |
||
|
|
13c8c768c7 |
Added the boot-at-EL1 option to the S32Z280 entry path (#690)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The Armv8-R AEM FVP entry path has carried TX_R52_BOOT_AT_EL1 since it was written, for the case its own comment describes: "an earlier boot stage or a vendor EL2 monitor has already dropped privilege to EL1". This board's entry path did not, so a kernel could not be built as a guest on it at all -- and it is the board where that matters most, because it is the one with silicon behind it. The bracket is the whole change. Everything from the Thumb reset trampoline to the ERET goes inside #ifndef TX_R52_BOOT_AT_EL1, and the #else supplies a one-instruction A32 _start that branches to el1_entry. A32 AND NOT T32, which is the one real difference from the standalone entry. The core resets in Thumb state here because the RTU boot instruction NXP plants is a T32 branch, but a guest is not reached by reset: it is reached by the monitor's ERET, and the monitor chooses the state through SPSR.T. Get the two out of agreement and the guest dies on its first instruction with an undefined-instruction exception, which looks exactly like a bad entry address and sends the reader to the loader instead of to the ERET. WHAT THE MONITOR INHERITS is enumerated at the #ifndef, next to the code it replaces rather than in a document, because that is where somebody adding a third board will be looking. This board's EL2 block is considerably larger than the model's, and each item on the list is something a guest at EL1 provably cannot do rather than something it merely does not: CNTFRQ is writable only at the highest implemented exception level and reads zero out of reset; HCPTR.TCP10/TCP11 reset set, trapping every EL1 floating-point access; HSCTLR.TE is an EL2 register (SCTLR.TE is EL1's, and el1_entry still clears it below); ICC_HSRE.SRE makes every other ICC_* and ICH_* register exist at all; the low-latency peripheral port enables reset to zero and an EL1 write to that register traps to EL2; and the TCM enables are per-core with ENABLEEL2 SILENTLY IGNORED from EL1 -- measured on both BTCM and CTCM, the base took and bit 0 took while bit 1 stayed clear. CNTHCTL.PL1PCTEN and PL1PCEN are the deliberate omission from that list, and the note says why. This path opens both, because a standalone kernel owns the physical timer. A monitor that TIME-partitions its guests must not: a partition's physical time keeps running while it is descheduled, so a guest reading it can observe that it was not running. That is the monitor's decision rather than this file's, which is why the list says what a guest cannot do rather than what a monitor should. Verified both ways. The three standalone images build and link unchanged, and a kernel built with the option boots at EL1 on a S32Z280-594EVB under an EL2 monitor, runs two threads through a queue and a semaphore, and reports back -- with no other change to the kernel or to its port. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
8c681c188e |
Asserted that the Cortex-R52 port refuses the options it documents as refused (#687)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
ports/cortex_r52/gnu/CMakeLists.txt rejects two option combinations at
configure time -- TX_R52_ENABLE_VFP without TX_R52_FLOAT_ABI=hard, and
TX_R52_ENABLE_FIQ_NESTING without TX_R52_ENABLE_FIQ. Nothing exercised
either. A guard that has stopped firing looks exactly like a guard nobody
has tripped, so both could have been silently disabled by a typo in a
variable name at any point and no check would have noticed.
check_gcc.sh gains a sixth stage that configures each rejected combination
and asserts the configure fails. It needs no new harness: the script already
runs cmake as a subprocess for the CMake example builds, and the port's
CMakeLists is included by the toolchain file alone, so no example flags are
needed and the two negative configures stop almost immediately.
Two things the stage does that a thinner version would not.
It asserts the message text, not just the exit status. A configure that
fails for an unrelated reason would otherwise be recorded as a guard doing
its job.
It also configures the SUPPORTED combination and requires that to succeed.
Two negative assertions on their own are satisfied by a guard that rejects
everything -- the port would be unbuildable and the check would still pass.
The positive case is what separates a guard that works from one that is
merely always on.
Verified negatively, four deliberate breaks, each caught:
VFP guard condition forced false -> "configure succeeded, but this
combination cannot build"
FIQ nesting guard forced false -> the same, on that case
guard fires, message text changed -> "configure failed, but not on the
expected guard"
VFP guard condition forced true -> "the supported combination was
refused"
The fourth is the one the positive case exists for and the only one a
refusals-only stage would have missed. ports/cortex_r52/gnu/CMakeLists.txt
was confirmed byte-identical to dev afterwards.
Full run passes: exit 0, all six stages, and --asm-only correctly skips the
new one.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
57390fc0fe |
Refused the Cortex-R52 VFP option without a hard float ABI (#686)
TX_R52_ENABLE_VFP with the default soft float ABI cannot build. The option defines TX_ENABLE_VFP_SUPPORT, which enables the VMRS, VSTMDB and VLDMIA blocks in the port assembly, and -mfloat-abi=soft leaves the assembler with no FPU to accept them. The configuration fails with eight errors of the form tx_thread_system_return.S:118: Error: selected processor does not support `vmrs r4,FPSCR' in ARM mode none of which mentions the float ABI, so a user has to reason from VMRS back to the option that enabled it. The guard that exists said otherwise. It warned that "the compiler will not emit floating-point instructions, so the VFP context path will never be exercised", which describes a build that succeeds and is merely pointless -- and then let configure finish, so the warning scrolled past well before the assembler errors appeared. It is now a FATAL_ERROR that names the fix, which is what the same file already does six lines above for TX_R52_ENABLE_FIQ_NESTING without TX_R52_ENABLE_FIQ. That combination is rejected for being "meaningless", while this one, which cannot assemble at all, was only warned about. The severities were the wrong way round. The option's definition also moves below the check, so the block reads like the FIQ nesting one. The ABI is not promoted to hard automatically. TX_R52_FLOAT_ABI is a cache variable the user may have set deliberately, and silently overriding an explicit choice is worse than refusing a combination that cannot work. readme_threadx.txt carried the same claim, and its option list marked the FIQ nesting dependency inline but not this one. Both corrected. No regression test. Nothing in the tree asserts a configure-time failure -- there is no harness for it, and the sibling FIQ nesting guard has none either -- so a test for this would have to introduce that mechanism for one case. Verified by hand in both directions instead: the soft-ABI combination now stops at configure with the message above, and the hard-ABI feature build (VFP, FIQ, IRQ nesting, FIQ nesting) builds its eight images clean and passes ctest 8/8 on the Armv8-R AEM FVP. scripts/check_gcc.sh passes unchanged; it configures the Cortex-R52 CMake stage without TX_R52_ENABLE_VFP, so the new branch is not on its path. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
a2800fef16 |
Covered the SMP suspension teardown and the long byte pool search (#677)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
Against the merged SMP coverage report -- every build configuration instrumented and unioned -- sixty-four lines of common_smp/src were uncovered, 5114 of 5178. Fifty-three of them are closed here and the report reads 5167 of 5178. The SMP coverage floor goes from 98 to 99 with it. Thirty-six of the sixty-four were one loop repeated four times: the walk in tx_block_pool_delete, tx_byte_pool_delete, tx_event_flags_delete and tx_queue_delete that releases every thread suspended on the object with TX_DELETED. The suite deletes all four object types after every single test, and that is exactly why the loop never ran. test_control_cleanup in the ThreadX suite deletes the application's objects first and its threads last, so a test that ends with a thread parked on a queue has that thread walked out of it by tx_queue_delete. The SMP suite's cleanup deletes the threads first, deliberately -- it was changed so that no application-owned object is still referenced when the object loops run, which is what stopped a class of teardown hang. The side effect is that all four deletes now run against an empty suspension list. tx_semaphore_delete is the one member of the family that was already covered, because threadx_semaphore_delete_test deletes a busy semaphore on purpose. threadx_object_delete_suspension_test is the same idea for the other four. Two threads suspend on each of a block pool, a byte pool, an event flags group and a queue; the control thread waits on the object's own suspended count through tx_*_info_get rather than on an ordering it cannot guarantee across four cores, deletes the object, and checks both waiters came out with TX_DELETED. Two waiters rather than one so the loop takes its back edge as well as its body, and every wait is bounded in ticks so a suspension that never arrives fails the test instead of hanging it. threadx_trace_entry_update_test and threadx_thread_misaligned_stack_test are ports of the two tests that closed the equivalent gaps in common/src, and close fourteen more lines here: tx_block_allocate 123, 175, 182, 319 and 326, tx_byte_allocate 130, 210, 217, 359 and 366, tx_thread_system_suspend 504 and 560, tx_trace_object_register 221, and tx_thread_create 133. The one substantive change is core confinement. The trace test needs thread 0 to suspend and thread 1 to then release what it waits for; on four cores thread 1 gives the block back before thread 0 has suspended and the update block behind the suspension is never reached, so both threads are excluded from cores 1 to 3. The misaligned stack test needed no such change. threadx_byte_memory_long_search_test closes three of the eleven in tx_byte_pool_search. Lines 264, 267 and 270 are the TX_BYTE_POOL_MULTIPLE_BLOCK_SEARCH limit -- twenty on this port -- where a long search drops and retakes protection so that it cannot lock the other cores out for the whole walk. No byte pool in the suite ever had twenty fragments. This one is filled with small chunks until it refuses another and then has every second chunk released, so the free fragments are never adjacent and cannot be merged, and the request is larger than any of them but smaller than the pool's theoretical total, which is what makes _tx_byte_pool_search walk rather than refuse at the door. The layout is asserted rather than assumed: the test checks the fragment count and checks the probe request really does fail before the workers start, because either would otherwise turn it into a silent no-op. Eleven lines remain and they are not a to-do list. Eight are the delay loop in tx_byte_pool_search that fires when another thread claims the pool inside the window the search opens. The Linux SMP port serialises all four cores on one pthread mutex, so that window is an unlock immediately followed by a lock on that mutex, and glibc hands an uncontended mutex straight back to the thread that just released it: measured over 180,003 windows across three cores, with zero handovers. The shipped test therefore does 250 searches per worker rather than the sixty thousand that probe used, because the twenty-block threshold is crossed by the first search. The other three are in tx_thread_smp_utilities. Line 149 is a range guard placed after the shift it is meant to guard, so reaching it needs a shift by the width of the type; the fix is to move the check above the shift, matching the TX_MAX_PRIORITIES > 32 variant of the same function, and that belongs in its own change. Lines 1073 and 1074 need a mutex owner that is genuinely executing on another core when a waiter suspends, and three shapes were tried without producing one on this port. Measured twice before and twice after, every gcda deleted between runs and 570 of 570 tests passing each time: 5114 of 5178 both times before, 5167 of 5178 both times after. Branch coverage goes from 2768 to 2821 and 2823 of 3548. A floor of 99 needs 5127, so the ratchet lands with forty lines of headroom against a numerator that has been seen moving by two between runs. Assisted-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
ff7fbe02f8 |
Added a GCC check for the Arm ports and ran it in CI (#675)
* Assembled the module ports, which no check had ever compiled
scripts/check_clang.sh globbed ports_module/*/gnu/src, which does not exist --
the module ports keep their assembly in module_manager/src. The [ -d ] guard
skipped it in silence, so 116 assembly files across nine Arm module ports were
assembled by no check, with either compiler, in the script whose own comments
state three times that "a port that is simply absent from the count reads as
covered". Stage 1 goes from 724 of 724 to 840 of 840; the feature-macro stage
had the same gap and goes from 412 files to 469.
Correcting the path exposed five defects, and only one of them was a build
failure. The other four assembled cleanly and did the wrong thing, because GAS
runs the C preprocessor on .S and not on .s:
ports_smp/cortex_a7_smp/gnu/src/tx_thread_smp_unprotect.s, the only .s in a
directory of twenty-one .S, ignored all four of its own feature macros. It
wrote the caller's LR into the protection structure on every unprotect -- a
store guarded by TX_MPCORE_DEBUG_ENABLE -- sent an unconditional SEV, and
returned through both BX lr and MOV pc, lr. Its cortex_a5_smp and
cortex_a9_smp siblings are .S.
ports_module/cortex_m33/.../tx_thread_stack_build.s emitted both arms of an
#ifdef TX_SINGLE_MODE_SECURE, so the non-secure LR value overwrote the secure
one and the secure build got the wrong frame.
ports_module/cortex_m23/.../tx_thread_context_{save,restore}.S carried the
POP {r0, lr} that check_clang.sh's own comment describes as the reason the
feature-macro stage exists. The 16-bit Thumb POP takes r0-r7 and pc only.
The identical fix already sits in ports/cortex_m23/gnu/src; the module copy
never got it because nothing scanned it.
ports_module/cortex_m23/.../tx_thread_secure_stack_initialize.S used MOV
rather than MOVS for an 8-bit immediate, latent behind TX_SINGLE_MODE_SECURE.
Both siblings in the same directory already use MOVS.
ports_module/cortex_a7/gnu/module_manager/src is the one that failed to
assemble, on GCC 14.3 as well as on LLVM: #define SYS_MODE was never
expanded, so #SYS_MODE reached the assembler as an undefined symbol.
Twenty-nine .s files under gnu trees are renamed to .S. Every one of them is
already named .S by the build scripts that compile it, so this repairs those
scripts rather than churning them -- ports_module/cortex_a7's build_threadx.bat
names all eighteen with a capital S, and works today only on a case-insensitive
filesystem. Renaming rather than converting the #defines to GNU assignments is
what fixes the #ifdef blocks as well as the constants; the assignments would
have fixed two files and left twenty-seven silently ignoring their macros.
Files with no preprocessor directives are left as .s: they are not broken, and
check_ports.sh gains a check that keeps them that way. Only the gnu trees are
checked there -- the IAR, Arm Compiler 5 and Keil assemblers preprocess .s
themselves, and about three hundred files in this repository rely on that.
Verified with both toolchains on the same tree: 840 of 840 assembled by
ATfE 22.1.0 and by arm-gnu-toolchain 14.3.rel1, all five stages of
check_clang.sh green, and check_ports.sh green including the reproducibility
check. The new check was shown to fail by planting a copy of the file it was
written for.
No regression test accompanies this. The assembly it covers is executed by no
host test, and the check itself going from 724 files to 840 is the coverage
AGENTS.md asks for -- together with the new check_ports.sh section, which is
what stops the class recurring.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Fixed the AArch64 samples, none of which had ever linked with GCC
Every AArch64 gnu example build failed at the sample link, all 27 of them --
13 under ports/ and 14 under ports_smp/:
libg.a(libc_a-init.o): in function `__libc_init_array':
undefined reference to `_init'
relocation truncated to fit: R_AARCH64_CALL26 against undefined
symbol `_init'
libg.a(libc_a-fini.o): in function `__libc_fini_array':
undefined reference to `_fini'
build_threadx_sample.sh links with -nostartfiles, which is correct for a port
carrying its own reset path, and that drops crti.o and crtn.o along with
everything else. startup.S calls __libc_init_array by design, and newlib's
implementation calls _init, which crti.o is what defines. The AArch32 scripts
are unaffected: they use nosys.specs and never reach __libc_init_array.
The fix links crti.o and crtn.o explicitly, bracketing the object list -- the
first must precede every .init contribution and the second must follow all of
them, so their position is load-bearing rather than stylistic. Both paths come
from the compiler's own -print-file-name, so nothing here hard-codes a
toolchain layout.
The atfe branch sets both to empty, deliberately: picolibc's __libc_init_array
does not call _init, those 27 images link today, and adding crti.o would change
a working link for no reason. That is also why check_clang.sh is green on these
and does not list them as expected to fail -- the LLVM path never reached the
gap, so nothing has ever linked them and failed.
Fixed in ports_arch/ARMv8-A/threadx/ports/gnu/example_build, which is the
single source for both the ports/ and ports_smp/ copies, then regenerated with
update.sh --port-sets tx,tx_smp. The 27 generated copies are in this commit
because ports_arch_check compares them.
Verified: all 27 link with arm-gnu-toolchain 14.3.rel1 aarch64-none-elf, where
0 of 27 did before; _init and _fini disassemble to the expected crti prologue
and crtn epilogue over a ret; check_clang.sh with ATfE 22.1.0 is still green on
all five stages, including the 42 script-driven example builds; check_ports.sh
is green including the reproducibility check.
No regression test: these are link-only example images that no host test
executes. What guards them is check_clang.sh's example stage today, and
check_gcc.sh's, which is the next change and is the reason this was found.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Added a GCC check for the Arm ports, which nothing had ever compiled
GCC is the project's declared default compiler (AGENTS.md, "The default
compiler for the project is GCC 14 on Linux"), it is what the gnu ports exist
for, and it is what nearly every downstream user builds with -- and nothing in
CI compiled a line of any port with it. The only cross-compilation check that
ran was the LLVM one, so the ATfE path was better guarded than the GNU one, on
ports whose directory is literally named gnu. ci_cortex_m covers four port
families; this covers forty.
Five stages, mirroring scripts/check_clang.sh stage for stage:
1. assemble every .S and .s of every Arm gnu port -- 840 files
2. assemble again the parts behind TX_ENABLE_VFP_SUPPORT,
TX_ENABLE_FIQ_SUPPORT, TX_LOW_POWER and
TX_ENABLE_EXECUTION_CHANGE_NOTIFY -- 469 files
3. compile common/src for one core per architecture profile -- 185 x 9
4. link the script-driven example builds -- 42
5. link the CMake-driven Cortex-R52 images -- 5
Two scripts rather than one with a --toolchain flag: the flag surface differs
(a prefixed driver against --target=), the C library differs, and the set of
examples that can link differs. Folding them together makes it easy to weaken
one check while working on the other.
Two toolchains, both required. Arm ships arm-none-eabi and aarch64-none-elf as
separate downloads and PORT_TARGET maps every port to one of exactly those two
triples, so --arm-none-eabi and --aarch64-none-elf each take a driver or the
directory holding it, defaulting to the environment and then to PATH. A missing
one is a hard error rather than a soft skip: letting a run cover half the tree
and still report "all checks passed" is the failure this script exists to end.
PORT_TARGET is copied verbatim from check_clang.sh, including its warning not
to prefix-match core names -- cortex_a5* also matches the AArch64 cortex_a53.
VFP_EXTRA is the one map that is not a copy, and check_clang.sh's comment about
it is false for GCC. That comment says the A-profile defaults are already
correct; arm-none-eabi-gcc defaults to -mfloat-abi=soft, which disables the FPU
outright, so every VFP file fails with "selected processor does not support
'vmrs r1,FPSCR' in ARM mode". -mfloat-abi=hard alone is the fix and is the
right one, because it selects the core's own default FPU rather than naming a
-d16 one -- which is the trap the clang script warns about, since the
A-profile paths save D16-D31. Cortex-R4 is the exception in both scripts and
for the same reason: its FPU is an option rather than part of the core, so an
explicit -mfpu is required. Every value was measured against 14.3.rel1.
Stage 4 *unsets* TOOLCHAIN rather than setting it. The example build scripts
already default to GNU, and a stray TOOLCHAIN=atfe from a developer's shell
would otherwise make this stage silently check the other compiler. It cleans
the example directories on both sides, because the success test is the
existence of sample_threadx.out rather than the driver's exit status, and a
stale image from a previous toolchain would report success. Failure logs are
printed unfiltered: a missing tool says "command not found", and GNU ld's
undefined-symbol lines carry no "error:" at all.
Every skip is printed by name with a reason, per the house rule check_clang.sh
states three times -- a port simply absent from the count reads as covered.
This script also says outright that arm9 and arm11 are Arm and are skipped for
having no PORT_TARGET entry, which the clang script's "not Arm" wording glosses.
Verified on this tree with arm-gnu-toolchain 14.3.rel1: all five stages green,
every count identical to check_clang.sh's on the same tree -- 840, 469, 185x9,
42, 5 -- in 4m28s.
The failure paths were tested, not assumed. A deliberately broken .S in a
module port is reported by name and line in stages 1 and 2 and exits 1, in
--quiet mode as well. Reverting the AArch64 _init/_fini fix on one port only
gives "FAIL: cortex_a53: example build produced no image", 41 of 42, and exit
1 -- and the log tail it prints contains no "error:" anywhere, which is why it
is not filtered. A missing or wrong toolchain path exits 1 naming which triple
was not found.
RISC-V is deliberately out of scope for this first version: both ports
assemble 8 of 8 with the project's own cmake flags, but adding them widens the
toolchain download and the review surface for a family that is not regressing.
No regression test accompanies this. The script is the test, it exercises no
runtime behaviour, and its own failure paths are exercised above.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Ran the GCC port check in CI, on dev as well as master
scripts/check_gcc.sh with nothing invoking it would be a script nobody runs.
This adds the workflow, modelled on clang_check.yml, and fixes a trigger gap in
that file at the same time.
One job, two cache steps. Arm ships AArch32 and AArch64 as separate downloads
and the script needs both, so two caches keep the checks list short and let a
single invocation see both compilers. The AArch32 cache path and key match
cortex_m's exactly, so the two workflows share one entry rather than each
holding its own copy of the same archive -- noted in a comment, because the
only symptom of breaking that is a slower run.
Both triggers name dev. A workflow that triggers only on master gates no pull
request anybody opens; that is the defect ports_arch_check.yml carries a
comment about, and it cost cortex_m three months of failing in seven seconds
unnoticed. push is included as well as pull_request so dev's own history has a
baseline and a bad squash-merge is caught rather than waiting for the next PR.
The checksum suffix is .sha256asc and it is not interchangeable with .sha256.
Arm publishes both for this release, and verified 26 Aug 2026, the .sha256 file
for arm-none-eabi contains a 32-character MD5 rather than a SHA-256, so
sha256sum -c on it fails with "no properly formatted checksum lines found".
.sha256asc is a plain sha256sum-format line for both triples. The plan warned
that this suffix had changed between releases; the sharper truth is that both
suffixes exist simultaneously and one of them is not a SHA-256 at all. Recorded
in a comment beside the step.
Verified before writing them in rather than copied: both archive URLs and both
checksum URLs resolve, the archives are xz, the checksum files are
sha256sum-format for .sha256asc, and the AArch64 archive extracts to
arm-gnu-toolchain-14.3.rel1-x86_64-aarch64-none-elf/bin/aarch64-none-elf-gcc,
which is the path the workflow builds.
The paths: lists are duplicated between push and pull_request rather than
shared through a YAML anchor, deliberately: GitHub Actions' parser does not
dependably honour anchors and the failure mode is the workflow refusing to
parse, which is the cortex_m failure again. Ten duplicated lines are cheaper.
clang_check.yml's paths: list was missing CMakeLists.txt, cmake/ and
common_smp/, so that check did not run when files it reads changed -- the
ports_smp example builds compile common_smp/src and its CMake stage reads the
toolchain file and the top-level project. Both lists are now identical apart
from each file's own name, and both say so.
cortex_m is kept rather than deleted, against the plan's recommendation. It
builds four ports *through CMake*, and that is the only thing exercising
cmake/cortex_m*.cmake and the top-level CMakeLists for the M profile; this
script's CMake stage covers cortex_r52 only. The overlap is the assembly and
the C sources, not the build system, so deleting it would lose coverage rather
than remove a duplicate. Said so in the workflow header.
The script is passed explicit toolchain paths rather than left to find the
drivers on PATH, so nothing about the runner image can decide which compiler
runs, and it prints both versions it resolved.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
5b94bad6a2 |
Fixed the AArch64 samples, none of which had ever linked with GCC (#673)
Every AArch64 gnu example build failed at the sample link, all 27 of them --
13 under ports/ and 14 under ports_smp/:
libg.a(libc_a-init.o): in function `__libc_init_array':
undefined reference to `_init'
relocation truncated to fit: R_AARCH64_CALL26 against undefined
symbol `_init'
libg.a(libc_a-fini.o): in function `__libc_fini_array':
undefined reference to `_fini'
build_threadx_sample.sh links with -nostartfiles, which is correct for a port
carrying its own reset path, and that drops crti.o and crtn.o along with
everything else. startup.S calls __libc_init_array by design, and newlib's
implementation calls _init, which crti.o is what defines. The AArch32 scripts
are unaffected: they use nosys.specs and never reach __libc_init_array.
The fix links crti.o and crtn.o explicitly, bracketing the object list -- the
first must precede every .init contribution and the second must follow all of
them, so their position is load-bearing rather than stylistic. Both paths come
from the compiler's own -print-file-name, so nothing here hard-codes a
toolchain layout.
The atfe branch sets both to empty, deliberately: picolibc's __libc_init_array
does not call _init, those 27 images link today, and adding crti.o would change
a working link for no reason. That is also why check_clang.sh is green on these
and does not list them as expected to fail -- the LLVM path never reached the
gap, so nothing has ever linked them and failed.
Fixed in ports_arch/ARMv8-A/threadx/ports/gnu/example_build, which is the
single source for both the ports/ and ports_smp/ copies, then regenerated with
update.sh --port-sets tx,tx_smp. The 27 generated copies are in this commit
because ports_arch_check compares them.
Verified: all 27 link with arm-gnu-toolchain 14.3.rel1 aarch64-none-elf, where
0 of 27 did before; _init and _fini disassemble to the expected crti prologue
and crtn epilogue over a ret; check_clang.sh with ATfE 22.1.0 is still green on
all five stages, including the 42 script-driven example builds; check_ports.sh
is green including the reproducibility check.
No regression test: these are link-only example images that no host test
executes. What guards them is check_clang.sh's example stage today, and
check_gcc.sh's, which is the next change and is the reason this was found.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
9c32abb17d |
Assembled the module ports, which no check had ever compiled (#672)
scripts/check_clang.sh globbed ports_module/*/gnu/src, which does not exist --
the module ports keep their assembly in module_manager/src. The [ -d ] guard
skipped it in silence, so 116 assembly files across nine Arm module ports were
assembled by no check, with either compiler, in the script whose own comments
state three times that "a port that is simply absent from the count reads as
covered". Stage 1 goes from 724 of 724 to 840 of 840; the feature-macro stage
had the same gap and goes from 412 files to 469.
Correcting the path exposed five defects, and only one of them was a build
failure. The other four assembled cleanly and did the wrong thing, because GAS
runs the C preprocessor on .S and not on .s:
ports_smp/cortex_a7_smp/gnu/src/tx_thread_smp_unprotect.s, the only .s in a
directory of twenty-one .S, ignored all four of its own feature macros. It
wrote the caller's LR into the protection structure on every unprotect -- a
store guarded by TX_MPCORE_DEBUG_ENABLE -- sent an unconditional SEV, and
returned through both BX lr and MOV pc, lr. Its cortex_a5_smp and
cortex_a9_smp siblings are .S.
ports_module/cortex_m33/.../tx_thread_stack_build.s emitted both arms of an
#ifdef TX_SINGLE_MODE_SECURE, so the non-secure LR value overwrote the secure
one and the secure build got the wrong frame.
ports_module/cortex_m23/.../tx_thread_context_{save,restore}.S carried the
POP {r0, lr} that check_clang.sh's own comment describes as the reason the
feature-macro stage exists. The 16-bit Thumb POP takes r0-r7 and pc only.
The identical fix already sits in ports/cortex_m23/gnu/src; the module copy
never got it because nothing scanned it.
ports_module/cortex_m23/.../tx_thread_secure_stack_initialize.S used MOV
rather than MOVS for an 8-bit immediate, latent behind TX_SINGLE_MODE_SECURE.
Both siblings in the same directory already use MOVS.
ports_module/cortex_a7/gnu/module_manager/src is the one that failed to
assemble, on GCC 14.3 as well as on LLVM: #define SYS_MODE was never
expanded, so #SYS_MODE reached the assembler as an undefined symbol.
Twenty-nine .s files under gnu trees are renamed to .S. Every one of them is
already named .S by the build scripts that compile it, so this repairs those
scripts rather than churning them -- ports_module/cortex_a7's build_threadx.bat
names all eighteen with a capital S, and works today only on a case-insensitive
filesystem. Renaming rather than converting the #defines to GNU assignments is
what fixes the #ifdef blocks as well as the constants; the assignments would
have fixed two files and left twenty-seven silently ignoring their macros.
Files with no preprocessor directives are left as .s: they are not broken, and
check_ports.sh gains a check that keeps them that way. Only the gnu trees are
checked there -- the IAR, Arm Compiler 5 and Keil assemblers preprocess .s
themselves, and about three hundred files in this repository rely on that.
Verified with both toolchains on the same tree: 840 of 840 assembled by
ATfE 22.1.0 and by arm-gnu-toolchain 14.3.rel1, all five stages of
check_clang.sh green, and check_ports.sh green including the reproducibility
check. The new check was shown to fail by planting a copy of the file it was
written for.
No regression test accompanies this. The assembly it covers is executed by no
host test, and the check itself going from 724 files to 840 is the coverage
AGENTS.md asks for -- together with the new check_ports.sh section, which is
what stops the class recurring.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
147754cc86 |
Enforced a coverage floor on the merged report (#667)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The coverage summary reported a percentage and could not fail. Coverage could fall from 99.97% to anything at all and every check stayed green, against an AGENTS.md that asks for 100% test coverage -- a stated requirement measured with a gauge that had no failure mode. CodeCoverageSummary already takes thresholds and fail_below_min; neither was set. Both are now, through a new coverage_thresholds input on the template, because the two suites do not sit at the same figure: ThreadX 99, SMP 98. Three things were probed against the pinned action on a runner before picking those numbers, using the real merged.xml files from the dev push run of #666. The floor compares the line rate and nothing else. That mattered because branch coverage is around 78% in both suites while line coverage is 98.8-100%, so a floor aimed at the line figure would have been an immediate red wall had it tested branches or the lower of the two. The ThreadX report at 100.00% lines and 77.67% branches clears a floor of 99. The thresholds are whole numbers. '99.9 100' -- the value this was meant to be -- is rejected with 'System.ArgumentException - Threshold parameter set incorrectly.', and the step fails whether or not fail_below_min is set. So the choice is 99 or 100 with nothing between. 100 would fail on a race. tx_thread_system_resume.c:529 is reached by timing rather than by construction and flaps between runs of the same green tree, which is why #666 left it; 4502/4503 fails a floor of 100 and clears one of 99. A coverage gate that goes red on a coin toss is how coverage gates get switched off. SMP is 5114/5178 lines, 98.76%, with 64 uncovered lines across 11 files of common_smp/src -- #666 closed the equivalent gaps in common/src only. A shared floor of 99 would have failed that job on every run while ThreadX passed. One limit is recorded in the file rather than fixed: an empty report reads as 100%. gcovr writes line-rate="1.0" beside lines-valid="0" when it finds no data, and the action prints 'Line Rate = 100% (0 / 0)' and passes any floor. The check for that is the emptiness assertion #664 put in each suite's coverage.sh, not this one. Also corrected two stale filenames in the deploy job's comment: since #665 each coverage artifact carries merged.xml, not default_build_coverage.xml. Verified on the runner. Assisted-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f89d65f041 |
Covered the trace entry update paths and the misaligned stack adjustment (#666)
Against the merged coverage report -- every build configuration instrumented and
unioned -- eighteen lines of common/src were uncovered. Seventeen of them were
one missing scenario rather than eighteen separate gaps.
tx_block_allocate, tx_byte_allocate, tx_thread_system_suspend and
tx_thread_system_resume each carry blocks under TX_ENABLE_EVENT_TRACE that go
back and patch a trace entry once the call has done its work, all of the shape:
if (entry_ptr != TX_NULL)
{
if (time_stamp == entry_ptr -> tx_trace_buffer_entry_time_stamp)
entry_ptr comes from _tx_trace_buffer_current_ptr, which stays TX_NULL until
tx_trace_enable is called at run time. Building with TX_ENABLE_EVENT_TRACE is
not enough, and exactly one test in the suite enables tracing --
threadx_trace_basic_test -- which tests the enable API itself and never calls
either allocator. So those blocks sat in the report's denominator and never in
its covered set.
threadx_trace_entry_update_test enables tracing and then drives both allocators
twice each, once on the path that succeeds immediately and once through a
suspension that a second thread satisfies, since each allocator carries one
update block on either side. It then sleeps so that the last runnable thread
suspends with nothing ready to take over: tx_thread_system_suspend lines 345 and
351 are on the branch that sets _tx_thread_execute_ptr to TX_NULL, and the two
allocator suspensions never reach it because the other thread was always ready.
The same test closes tx_trace_object_register's TX_NULL name break by creating a
semaphore with no name. A semaphore and not a thread deliberately: for
TX_TRACE_OBJECT_TYPE_THREAD the register function dereferences the pointer it is
given to read the thread's priority, so that type needs a real TX_THREAD behind
it. threadx_trace_basic_test makes the equivalent call only under
ifndef TX_ENABLE_EVENT_TRACE, against the no-op stub.
threadx_thread_misaligned_stack_test covers the remaining line,
tx_thread_create.c:136, where a stack that does not begin on a ULONG boundary
costs a ULONG of size so that rounding the start up cannot run past the end of
the caller's buffer. Every other test hands tx_thread_create an aligned stack.
That line is compiled only under TX_ENABLE_STACK_CHECKING, so it is absent from
three of the five configurations' reports rather than uncovered in them, and it
was verified under stack_checking_build.
Measured on the merged report, all 480 tests passing and 5 of 5 configurations
green: 4485 of 4503 covered before, 4502 of 4503 after, denominator unchanged.
One line remains, tx_thread_system_resume.c:529, and it is the report's last
flapping line rather than a standing gap -- two clean runs of the same tree gave
4503 of 4503 and 4502 of 4503. Reaching it by construction was tried twice and
failed both times, so it is left alone here. It needs the preempt disable flag
and the system state both clear, and tx_thread_resume raises the preempt disable
flag before calling _tx_thread_system_resume, as do the put and send paths;
creating a higher priority auto-start thread from thread context does not raise
it but does not reach the check either, which a probe showed is executed only
during initialisation, with the system state at TX_INITIALIZE_IN_PROGRESS.
Assisted-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
3e85bbd431 |
Instrumented every build configuration and merged their coverage (#665)
Only default_build_coverage carried -fprofile-arcs, because the gate was the build type and it is the only one of five whose name ends in _coverage. The other four build and run all their tests and their coverage was discarded. That is not redundancy thrown away: each configuration selects a different set of TX_ feature macros, so the code the other four compile is absent from the denominator rather than uncovered in it. TX_COVERAGE instruments a build regardless of its name, defaulting to OFF so a single configuration built by hand behaves as before. coverage.sh gains a --merge mode that unions the per-configuration JSON tracefiles, and cmake_bootstrap.sh runs it after the test loop so a local run produces the same merged report CI reads. The template sets TX_COVERAGE for build and test, and coverage_name moves to the merged report. Measured on the ThreadX suite, all 480 tests passing: default_build_coverage 3827 valid 3827 covered disable_notify_callbacks_build 3767 3766 stack_checking_build 3857 3856 stack_checking_rand_fill_build 3862 3861 trace_build 4123 4108 merged 4503 4487 The denominator grows by 676 lines, 17.7%, and the figure moves from 99.97% to 99.64%. The second one is honest, and the drop is the point rather than a regression: the denominator now includes code the old report never counted. The union also contains a file the old report did not contain at all -- tx_thread_stack_error_handler.c compiles only under TX_ENABLE_STACK_CHECKING, so it was not listed at 0%, it was simply absent. 177 files becomes 178. Coverage collection moved out of test() and now runs after the test loop, one configuration at a time. gcov writes its intermediate gcov files into the directory gcovr is rooted at, and coverage.sh roots every configuration at the repository root so filenames come out repo-relative. Five concurrent gcovr processes therefore share one scratch directory and delete each other's output: the first full run of this change passed all 480 tests and produced no report for three of the five configurations. Measured both ways -- two gcovr rooted at the repository root fail concurrently and succeed in sequence. CI would not have caught it, because test_tx.sh sets CTEST_PARALLEL_LEVEL=1 and takes the serial branch. Per-configuration output moved under coverage_report/per_configuration/ and is excluded from the Pages artifact. The deploy job merges the ThreadX and SMP artifacts into one tree and every configuration directory has the same name in both, so left at the top level one suite's would overwrite the other's on the published site. On the SMP suite, an earlier run of this change saw trace_build fail threadx_smp_time_slice_test and then hang, which raised the question of whether -fprofile-arcs perturbs a timing-sensitive test. It does not. Sixteen runs settle it, and the decisive one is that threadx_smp_time_slice_test failed ERROR #31 -- twice in a row under --repeat until-pass:2 -- on an uninstrumented build, in the exact shape CI runs, while three instrumented runs of that shape passed 5 of 5. In the CI shape, CTEST_PARALLEL_LEVEL=1 run.sh test all: TX_COVERAGE=OFF 3 runs 2 green, one ERROR #31 310 s TX_COVERAGE=ON 3 runs 3 green, 5/5 each 325-329 s So the test is a pre-existing flake on dev and instrumenting all five costs about 5% of the suite's wall clock. Separately, and also in both instrumented and uninstrumented builds, run.sh's parallel branch -- what a developer gets typing run.sh test all with no CTEST_PARALLEL_LEVEL -- hangs under its own load, four times in twelve runs. Several SMP tests create 1024 ThreadX threads by construction and the Linux port backs each with a pthread, so five configurations at once put on the order of 5000 threads on the machine. CI sets CTEST_PARALLEL_LEVEL=1 and does not take that branch. Assisted-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
b756220c43 |
Fixed the coverage report's paths and scoping (#664)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The Cobertura XML embedded absolute machine paths, and the flag that looked like it scoped the report to one build configuration was doing nothing at all. Both coverage.sh scripts had the defect; both are fixed here, because the SMP report is published to the same Pages site as the ThreadX one. Paths. -r was the build directory and -f pointed outside it, so gcovr could not express the sources relative to the root and fell back to absolute paths. The result named files as /home/runner/work/threadx/threadx/common/src/... while the <source> element beside them said build/default_build_coverage, so the two halves of the same file disagreed and nothing could map coverage back to the repository. -r is the repository root now and -f an absolute path beneath it, which gives filename="common/src/tx_block_allocate.c". Both must be absolute: -r ../../.. -f common/src produces a report of zero files and exits 0, which is the worst failure mode available here. Scoping. --object-directory does not restrict which gcda files are found -- it tells gcovr how to get from a gcda file back to the compiler's working directory. Pointed at an empty directory it still produced the full 177-file report. That was harmless only by accident, because -r build/$1 constrained the search instead; moving -r to the repository root removes that accident, so the two changes have to land together. Measured, with a second instrumented configuration deliberately made sparser than the first: scoped by the positional search path 3221 of 3827 lines -- the truth no search path, -r at the repo root 3827 of 3827 -- silently merged --object-directory at the sparse tree 3827 of 3827 -- scopes nothing So the search path is load-bearing, and it matters ahead of instrumenting all five configurations: without it each configuration would have reported the union as its own. An empty report is not an error to gcovr -- it warns and exits 0 -- and it carries line-rate="1.0" next to lines-valid="0", so a consumer reads no data at all as fully covered. No coverage threshold can catch that, since an empty report passes any threshold. Hence the explicit assertion that the report has content, which fires with exit 1 on an object directory that exists but is empty, where the old shape returned 177 files and exit 0. Also says out loud that ports/linux/gnu/src is deliberately outside the filter. gcno files exist for it and it is dropped without a word today. Number-neutral, and that was the test. Over the same frozen gcda, changing only the gcovr invocation: ThreadX 3827 of 3827 lines and 1993 of 1994 branches across 177 files, SMP 4739 of 4791 and 2417 of 2430 across 185, before and after alike. Same answers on gcovr 7.0, 8.3 and 8.6, so the change is not wedged to the current pin. End to end through run.sh, 96 of 96 and 110 of 110 pass with the reports written. Assisted-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
2d9b9f7417 |
Bumped gcovr off the 4.1 pin it had been held on since 2018 (#663)
The coverage tooling was pinned to gcovr 4.1, released in 2018, and that version is missing the two options the coverage work needs next: --json and --add-tracefile, which is how the five build configurations get merged into one report. This moves the pin to 8.6, the current release. The pin stays exact, and it stays hand-moved: it lives in a shell script, and no Dependabot ecosystem can parse that. Isolated deliberately, so that a movement in the coverage number caused by the tool could not be confused with one caused by a later change. Measured on the default_build_coverage tree of test/tx, over the same gcda with the same gcov, varying only the gcovr version: gcovr lines-valid branches-valid files 4.1 3827 1994 177 7.0 3827 1994 177 8.3 3827 1994 177 8.6 3827 1994 177 So the denominator does not move with the tool at all, and this bump moves no number. The plan this came from expected 3822 to become 3827; that figure does not reproduce, under gcc-13 or gcc-14, with or without --object-directory. The only variant that changes the count is dropping the -f filter, which collapses the report to nothing. Two things found while measuring, both recorded because they matter to what comes next. The coverage numerator is not deterministic. On an identical tree with an identical compiler, three consecutive runs of the full suite -- all 96 tests passing every time -- reported 3826, 3827 and 3827 covered lines. The line that flickers is tx_thread_system_resume.c:529, the preemption path of _tx_thread_system_resume, and it takes its guarding branch with it. It has been described as never executed; it is executed on some runs and not others. A coverage floor has to be set with that in mind, and the honest fix is a test that takes the path deliberately. Reading gcc-13 output, the compiler the runners actually use, gcovr 8.6 runs the existing coverage.sh unchanged: Cobertura XML and 181 HTML files, same 177 classes. --xml-pretty and --object-directory still work on 8.6 but are now deprecated aliases for --cobertura-pretty and --gcov-object-directory, worth knowing for whoever removes --object-directory next. Assisted-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
b6a00a2014 |
Added the Dependabot configuration the pinned actions need (#662)
The action references were pinned to commit SHAs in #660, and a SHA pin with nothing moving it is worse than a floating tag -- it holds CI on whatever was current the day it was written. That is exactly how actions/cache@v1 stayed in ci_cortex_m.yml until GitHub began auto-failing every request that used it. The drift measured before that catch-up: download-artifact four majors behind, checkout and upload-artifact three each, cache and upload-pages-artifact two, with nothing ever reporting it. This closes the loop, and the reference to .github/dependabot.yml that #660 left in each workflow's pinning comment. Weekly, github-actions only. Patch and minor are grouped into one pull request because they are the routine traffic and a queue reviewed one item at a time is a queue that gets ignored. Majors stay ungrouped, one each, because every breaking change this repository has met in an action has been a major. Two choices worth stating rather than leaving to be rediscovered. target-branch is dev. Dependabot reads this file from the default branch, which is master, but master is deliberately kept behind dev and pull requests belong where the regression suites gate them. The consequence is that landing this on dev arms it without firing it: nothing happens until a release merge carries the file to master. Setting target-branch also opts out of Dependabot security updates, which only run against the default branch -- a small cost for this ecosystem, since an action advisory arrives as an ordinary bump on the weekly run, but a real one. The pull-request limit is raised from the default five to ten. Nine actions are in use, and five would hold majors back with nothing saying that it had. No other ecosystem is configured, deliberately: external dependencies are forbidden, there are no submodules, and the one pinned tool -- gcovr in scripts/install.sh -- lives in a shell script no ecosystem can parse, so that pin keeps moving by hand. No sibling eclipse-threadx repository has a Dependabot configuration, so this sets the pattern rather than following one. The dependencies label it uses already exists here. Assisted-by: Claude Opus 5 <noreply@anthropic.com> |