Commit Graph
39 Commits
Author SHA1 Message Date
Frédéric Desbiens f89d65f041 Covered the trace entry update paths and the misaligned stack adjustment (#666)
Against the merged coverage report -- every build configuration instrumented and
unioned -- eighteen lines of common/src were uncovered. Seventeen of them were
one missing scenario rather than eighteen separate gaps.

tx_block_allocate, tx_byte_allocate, tx_thread_system_suspend and
tx_thread_system_resume each carry blocks under TX_ENABLE_EVENT_TRACE that go
back and patch a trace entry once the call has done its work, all of the shape:

    if (entry_ptr != TX_NULL)
    {
        if (time_stamp == entry_ptr -> tx_trace_buffer_entry_time_stamp)

entry_ptr comes from _tx_trace_buffer_current_ptr, which stays TX_NULL until
tx_trace_enable is called at run time. Building with TX_ENABLE_EVENT_TRACE is
not enough, and exactly one test in the suite enables tracing --
threadx_trace_basic_test -- which tests the enable API itself and never calls
either allocator. So those blocks sat in the report's denominator and never in
its covered set.

threadx_trace_entry_update_test enables tracing and then drives both allocators
twice each, once on the path that succeeds immediately and once through a
suspension that a second thread satisfies, since each allocator carries one
update block on either side. It then sleeps so that the last runnable thread
suspends with nothing ready to take over: tx_thread_system_suspend lines 345 and
351 are on the branch that sets _tx_thread_execute_ptr to TX_NULL, and the two
allocator suspensions never reach it because the other thread was always ready.

The same test closes tx_trace_object_register's TX_NULL name break by creating a
semaphore with no name. A semaphore and not a thread deliberately: for
TX_TRACE_OBJECT_TYPE_THREAD the register function dereferences the pointer it is
given to read the thread's priority, so that type needs a real TX_THREAD behind
it. threadx_trace_basic_test makes the equivalent call only under
ifndef TX_ENABLE_EVENT_TRACE, against the no-op stub.

threadx_thread_misaligned_stack_test covers the remaining line,
tx_thread_create.c:136, where a stack that does not begin on a ULONG boundary
costs a ULONG of size so that rounding the start up cannot run past the end of
the caller's buffer. Every other test hands tx_thread_create an aligned stack.
That line is compiled only under TX_ENABLE_STACK_CHECKING, so it is absent from
three of the five configurations' reports rather than uncovered in them, and it
was verified under stack_checking_build.

Measured on the merged report, all 480 tests passing and 5 of 5 configurations
green: 4485 of 4503 covered before, 4502 of 4503 after, denominator unchanged.

One line remains, tx_thread_system_resume.c:529, and it is the report's last
flapping line rather than a standing gap -- two clean runs of the same tree gave
4503 of 4503 and 4502 of 4503. Reaching it by construction was tried twice and
failed both times, so it is left alone here. It needs the preempt disable flag
and the system state both clear, and tx_thread_resume raises the preempt disable
flag before calling _tx_thread_system_resume, as do the put and send paths;
creating a higher priority auto-start thread from thread context does not raise
it but does not reach the check either, which a probe showed is executed only
during initialisation, with the system state at TX_INITIALIZE_IN_PROGRESS.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 08:14:38 -04:00
Frédéric Desbiens 3e85bbd431 Instrumented every build configuration and merged their coverage (#665)
Only default_build_coverage carried -fprofile-arcs, because the gate was the
build type and it is the only one of five whose name ends in _coverage. The
other four build and run all their tests and their coverage was discarded. That
is not redundancy thrown away: each configuration selects a different set of TX_
feature macros, so the code the other four compile is absent from the
denominator rather than uncovered in it.

TX_COVERAGE instruments a build regardless of its name, defaulting to OFF so a
single configuration built by hand behaves as before. coverage.sh gains a
--merge mode that unions the per-configuration JSON tracefiles, and
cmake_bootstrap.sh runs it after the test loop so a local run produces the same
merged report CI reads. The template sets TX_COVERAGE for build and test, and
coverage_name moves to the merged report.

Measured on the ThreadX suite, all 480 tests passing:

  default_build_coverage           3827 valid   3827 covered
  disable_notify_callbacks_build   3767         3766
  stack_checking_build             3857         3856
  stack_checking_rand_fill_build   3862         3861
  trace_build                      4123         4108
  merged                           4503         4487

The denominator grows by 676 lines, 17.7%, and the figure moves from 99.97% to
99.64%. The second one is honest, and the drop is the point rather than a
regression: the denominator now includes code the old report never counted. The
union also contains a file the old report did not contain at all --
tx_thread_stack_error_handler.c compiles only under TX_ENABLE_STACK_CHECKING, so
it was not listed at 0%, it was simply absent. 177 files becomes 178.

Coverage collection moved out of test() and now runs after the test loop, one
configuration at a time. gcov writes its intermediate gcov files into the
directory gcovr is rooted at, and coverage.sh roots every configuration at the
repository root so filenames come out repo-relative. Five concurrent gcovr
processes therefore share one scratch directory and delete each other's output:
the first full run of this change passed all 480 tests and produced no report
for three of the five configurations. Measured both ways -- two gcovr rooted at
the repository root fail concurrently and succeed in sequence. CI would not
have caught it, because test_tx.sh sets CTEST_PARALLEL_LEVEL=1 and takes the
serial branch.

Per-configuration output moved under coverage_report/per_configuration/ and is
excluded from the Pages artifact. The deploy job merges the ThreadX and SMP
artifacts into one tree and every configuration directory has the same name in
both, so left at the top level one suite's would overwrite the other's on the
published site.

On the SMP suite, an earlier run of this change saw trace_build fail
threadx_smp_time_slice_test and then hang, which raised the question of whether
-fprofile-arcs perturbs a timing-sensitive test. It does not. Sixteen runs
settle it, and the decisive one is that threadx_smp_time_slice_test failed
ERROR #31 -- twice in a row under --repeat until-pass:2 -- on an uninstrumented
build, in the exact shape CI runs, while three instrumented runs of that shape
passed 5 of 5. In the CI shape, CTEST_PARALLEL_LEVEL=1 run.sh test all:

  TX_COVERAGE=OFF   3 runs   2 green, one ERROR #31        310 s
  TX_COVERAGE=ON    3 runs   3 green, 5/5 each             325-329 s

So the test is a pre-existing flake on dev and instrumenting all five costs
about 5% of the suite's wall clock. Separately, and also in both instrumented
and uninstrumented builds, run.sh's parallel branch -- what a developer gets
typing run.sh test all with no CTEST_PARALLEL_LEVEL -- hangs under its own load,
four times in twelve runs. Several SMP tests create 1024 ThreadX threads by
construction and the Linux port backs each with a pthread, so five
configurations at once put on the order of 5000 threads on the machine. CI sets
CTEST_PARALLEL_LEVEL=1 and does not take that branch.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 08:13:33 -04:00
Frédéric Desbiens b756220c43 Fixed the coverage report's paths and scoping (#664)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The Cobertura XML embedded absolute machine paths, and the flag that looked
like it scoped the report to one build configuration was doing nothing at all.
Both coverage.sh scripts had the defect; both are fixed here, because the SMP
report is published to the same Pages site as the ThreadX one.

Paths. -r was the build directory and -f pointed outside it, so gcovr could not
express the sources relative to the root and fell back to absolute paths. The
result named files as /home/runner/work/threadx/threadx/common/src/... while
the <source> element beside them said build/default_build_coverage, so the two
halves of the same file disagreed and nothing could map coverage back to the
repository. -r is the repository root now and -f an absolute path beneath it,
which gives filename="common/src/tx_block_allocate.c".

Both must be absolute: -r ../../.. -f common/src produces a report of zero
files and exits 0, which is the worst failure mode available here.

Scoping. --object-directory does not restrict which gcda files are found -- it
tells gcovr how to get from a gcda file back to the compiler's working
directory. Pointed at an empty directory it still produced the full 177-file
report. That was harmless only by accident, because -r build/$1 constrained the
search instead; moving -r to the repository root removes that accident, so the
two changes have to land together. Measured, with a second instrumented
configuration deliberately made sparser than the first:

  scoped by the positional search path   3221 of 3827 lines -- the truth
  no search path, -r at the repo root    3827 of 3827 -- silently merged
  --object-directory at the sparse tree  3827 of 3827 -- scopes nothing

So the search path is load-bearing, and it matters ahead of instrumenting all
five configurations: without it each configuration would have reported the
union as its own.

An empty report is not an error to gcovr -- it warns and exits 0 -- and it
carries line-rate="1.0" next to lines-valid="0", so a consumer reads no data at
all as fully covered. No coverage threshold can catch that, since an empty
report passes any threshold. Hence the explicit assertion that the report has
content, which fires with exit 1 on an object directory that exists but is
empty, where the old shape returned 177 files and exit 0.

Also says out loud that ports/linux/gnu/src is deliberately outside the filter.
gcno files exist for it and it is dropped without a word today.

Number-neutral, and that was the test. Over the same frozen gcda, changing only
the gcovr invocation: ThreadX 3827 of 3827 lines and 1993 of 1994 branches
across 177 files, SMP 4739 of 4791 and 2417 of 2430 across 185, before and
after alike. Same answers on gcovr 7.0, 8.3 and 8.6, so the change is not
wedged to the current pin. End to end through run.sh, 96 of 96 and 110 of 110
pass with the reports written.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 17:16:51 -04:00
Frédéric Desbiens eabdb86409 Matched gcov to the compiler that produced the coverage data (#658)
gcov reads a data format tied to the compiler that produced it. coverage.sh
took whatever gcov was first on PATH, which was fine while the compiler was
also whatever was first on PATH. #656 made cmake/linux.cmake honour CC and
#657 made a compiler switch actually reconfigure the build, so that
assumption no longer holds, and the first person to use the new capability
would have hit this.

Measured on dev with both of those merged:

    CC=gcc-14 ./run.sh build default_build_coverage    # succeeds
    ./coverage.sh default_build_coverage               # exit 64

gcov says why, if asked directly:

    tx_block_allocate.c.gcno:version 'B42*', prefer 'B33*'

gcovr turns that into "GCOV returncode was 3" and exits 64 through a Python
traceback, after the tests have already passed. It reads like a coverage bug
rather than a toolchain mismatch, which is the part that would have cost
someone an afternoon.

gcov is now derived from CC rather than found on PATH, so the caller sets
one variable instead of remembering two. GCOV still overrides, for a
toolchain that does not follow the gcc/gcov naming, and a derived gcov that
does not exist is reported as such instead of surfacing as a traceback.

Verified, tx and smp, before and after:

    CC=gcc-14    was exit 64, now exit 0, 177 files and 1527/3827 lines
    CC unset     exit 0, 177 files and 1527/3827 lines, unchanged
    CC=gcc-99    exit 1 naming gcov-99 and CC, rather than a traceback
    GCOV=gcov-14 with CC=gcc-99, exit 0, so the override still wins

A mismatched pairing still fails, deliberately: reading a gcc-14 tree with
the default gcc-13 gcov is exit 64 as before. Producing a number from
mismatched data would be worse than refusing.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 15:55:18 -04:00
Frédéric DesbiensandClaude Opus 5 e46b1b0787 Asked the wait abort test for three windows, and failed a run that reached none (#649)
Four CI runs of the same tree, twenty configuration-runs in total, show this
test's budget being reached far more often than the first green run suggested,
and a pass being reported every time it was:

    trace_build            3 of 10 windows in 121 seconds
    disable_notify         3 of 10 windows in 121 seconds
    default_coverage       4 of 10 windows in 121 seconds
    stack_checking         7 of 10 windows in 121 seconds
    trace_build            0 of 10 windows in 121 seconds
    disable_notify         7 of 10 windows in 121 seconds
    stack_checking         3 of 10 windows in 121 seconds

Seven of twenty, and the shortfall message only ever reaches an artifact:
ctest is run with --output-on-failure, so a passing test's output is not in the
job log at all. The suite has been quietly losing most of this test's coverage
in whole configurations and reporting green.

The loop runs in two modes, not one. A window arrives in milliseconds in the
fast mode, and costs between 17 and 40 seconds in the slow one, with nothing in
between across those twenty runs. Ten windows are therefore unreachable inside
any budget worth having: at 40 seconds each that is 400 seconds, and the
unbounded runs measured before any of this took up to 726. Raising the budget
to cover the slow mode would trade a quiet loss of coverage for five
configurations approaching the sixty minute step timeout.

So ask for what a run can reach. Three windows cost 51 to 120 seconds in the
slow mode and under a second in the fast one, and the later hits repeat what
the first ones establish, so what is given up is small. The budget goes to 180
seconds because three windows at the worst rate measured is exactly the 120 it
was, which would have truncated at two.

The count is printed on every run rather than only on a short one. A number
that appears only on shortfall cannot be told apart from a number nobody
recorded.

Reaching the window no times at all is a different matter, and was the worst of
the seven. The check after the loop compares semaphore bookkeeping that a
window has to have touched to mean anything, so a run that reached none of them
compares a counter against the value it was initialised to and reports a pass
having verified nothing. That run now keeps trying to a 300 second ceiling, and
fails if it still has not reached the window. A genuine resonance that holds
for five minutes is worth a failure; the old behaviour was worth nothing.

The SMP copy keeps its count of twenty. It reaches them in under half a second
in all five of its configurations, in all four runs, so the slow mode has never
been observed there and the coverage is free. Both copies get the ceiling and
the unconditional report, so the logic stays identical between them.

Verified locally on all five configurations: the test reaches 3 of 3 in 5 to 14
seconds, and the full suites pass 96 of 96 and 110 of 110 run one test at a
time. With the handler's window made unreachable and the ceiling lowered to 5
seconds, the test stops after 6 seconds, prints the count it reached, and
reports ERROR #8 with the harness recording a failure rather than a pass. With
the count raised past what the budget allows, a run that reaches two windows
still passes, so falling short and reaching nothing stay distinct. The
TX_NOT_INTERRUPTABLE branch, which no configuration in either suite builds, was
compile-checked in both copies with the configurations' own compile commands.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 11:11:31 -04:00
Frédéric DesbiensandClaude Opus 5 fe40079353 Stopped the thread priority change test leaving core 0 to a finished thread (#647)
The SMP suite has been failing on threadx_thread_priority_change since 30 June.
The test reports SUCCESS and the process then never exits, so ctest kills it at
the thousand second timeout, twice, and the log carries nothing past the result
line. Instrumenting the harness teardown produced the state at the hang:

  last stage reached:         test_control_return: control thread resume returned
  _tx_thread_preempt_disable: 0
  core 0: current=thread 0 execute=thread 0
  thread test control thread  state=0 priority=0 threshold=0 core_control=1
  thread thread 0             state=1 priority=0 threshold=0 inherit=0

Core 0 is held by a thread in state 1, TX_COMPLETED, while the control thread
sits in state 0, TX_READY, at priority 0. The Linux SMP scheduler fills a core
only when _tx_thread_current_ptr for it is null, and clears that pointer only
for a thread carrying a deferred preemption, which thread 0 is not. So core 0
can never be handed on, and with TX_THREAD_SMP_ONLY_CORE_0_DEFAULT and
TX_SMP_NOT_POSSIBLE the control thread has nowhere else to go. The scheduler
re-reads the same state every two milliseconds for as long as it is allowed to.

What put thread 0 at priority 0 is the last thing this test does:

    thread_0.tx_thread_inherit_priority = 0;
    _tx_thread_smp_simple_priority_change(&thread_0, 16);

with the stated intent of reaching the branch where the new priority is below
the inheritance priority. Zero cannot reach that branch, because 16 is not less
than 0. The other branch runs instead, and that branch assigns the inheritance
priority as the thread's priority while the code after it links the thread into
the list for the new priority regardless. Thread 0 therefore came away claiming
priority 0 while living in the priority 16 list.

Both halves of that hurt. Priority 0 ties with the control thread, so resuming
the control thread raised no preemption and left the execute pointer alone. The
mismatch between the recorded priority and the list the thread is linked into
then means that completing thread 0 removes it from a list it was never in,
leaving core 0 pointing at it for good.

Give the inheritance priority a value above the new one, which is what the
branch the comment names actually requires, and put it back to
TX_MAX_PRIORITIES afterwards so nothing downstream reasons about an
inheritance that is not there. Hold protection across the call as well: this is
an internal routine that expects it, and it was being called in the open.

Measured before and after by printing the thread's state at the point the test
reports success. With the inheritance priority at 0 it is priority 0 threshold
0, matching the hang above, on every run. With it above the new priority it is
priority 16 threshold 16, which agrees with the list the thread is linked into,
and resuming the control thread preempts core 0 the ordinary way.

All five SMP configurations pass 110 of 110 at the parallelism CI uses.

The comparison in _tx_thread_smp_simple_priority_change deserves a second look
on its own account. When its else branch runs, the thread's recorded priority
and the list it is linked into disagree by construction. Only this test is
known to reach that branch, by supplying an inheritance priority that cannot
arise in ordinary operation, so nothing here claims a defect in shipped paths.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 08:52:06 -04:00
Frédéric DesbiensandClaude Opus 5 83dfbc3534 Made a teardown hang in the SMP suite say where it stopped (#646)
The SMP regression suite times out in CI on threadx_thread_priority_change and
the log carries nothing that says why. The reason the log is empty is
mechanical: test_control_return() opens with fflush(stdout), and that is the
last flush before either the test finishes or it wedges. Everything printed
after it sits in stdout's block buffer, and ctest discards that buffer when it
kills the process at the timeout. So the log ends at the test's own result line
no matter where the process actually stopped.

That is enough to place the hang, if not to explain it. The failing runs print
"SUCCESS!" and then stop, which means every check in the test body ran and
passed, and the wedge is somewhere between that flush and exit(). The run of
30 June shows the same signature before any of this test's waits were bounded,
so the hang is not the unbounded wait removed earlier, and the message added
then for an exhausted cap never appears. Retrying tells us nothing new either:
the suite spends 1000 seconds per attempt, twice, to reproduce the same silent
timeout.

Record how far teardown gets, and bound it. A stage variable is updated at each
step from test_control_return() through test_control_cleanup() to exit(), and a
watchdog thread armed on entry to test_control_return() reports the last stage
reached, the per-core scheduler state, and every thread on the created list,
then exits 99.

The watchdog covers teardown and not the test body, because the test body has
no bounded runtime to hold it to. Several tests here wait on a probabilistic
interrupt window: threadx_thread_wait_abort_and_isr_test has been measured
between 0.34 and 439 seconds while passing. Teardown is a fixed amount of work
that takes milliseconds, so a bound on it cannot turn a slow pass into a
failure. The default is 60 seconds, which also means a wedged run now reports
in one minute rather than burning the 2000 seconds two 1000-second attempts
cost today.

The report is written with write() rather than printf() because a wedged thread
may be holding the stdio lock, and a watchdog that blocked on that lock would
reproduce the silent timeout it exists to replace. For the same reason it reads
the ThreadX globals directly and takes no kernel lock; the values may be torn,
which is acceptable for a post-mortem and cannot deadlock.

One walk in test_control_cleanup() is bounded as well. The loop that steps past
the timer thread and the control thread has no terminating condition of its own
and spins for good if _tx_thread_created_count and the created list ever
disagree, which is one of the shapes the timeout could be taking. It now
reports and stops instead.

Off by default in the sense that matters: stderr stays empty and stdout keeps
its buffering, so output is unchanged on a passing run.
TX_TEST_TEARDOWN_TIMEOUT overrides the bound in seconds and zero disables the
watchdog; TX_TEST_TEARDOWN_TRACE echoes each stage as it is reached and
line-buffers stdout so the surrounding output survives a kill too.

Verified against an injected hang at the point the failing runs stop: the
watchdog fires, names the stage, and exits 99. The dump is already informative,
showing thread 0 left at priority 0 with threshold 0 and inherit 0, the same
priority as the control thread, with core 0's execute pointer still on it. All
five SMP configurations pass 110 of 110 at the parallelism CI uses, in both
quiet and trace modes, and the suite runtime is unchanged.

Only the SMP harness is instrumented. The non-SMP suite has not shown this
hang.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 08:51:05 -04:00
Frédéric DesbiensandClaude Opus 5 a5483f0773 Bounded the wait for the delayed suspension window (#645)
threadx_thread_delayed_suspension_test waits for an interrupt to land while
thread 2 is part way through suspending, and waits for it with no bound:

    while(delayed_suspend_set == 0)
    {
        tx_thread_wait_abort(&thread_2);
        tx_thread_relinquish();
    }

How long that takes depends on the build to a degree that is easy to miss. The
loop finishes in between a tenth of a second and three seconds in four of the
five ThreadX configurations. In trace_build it took 490 seconds, which was 40
percent of the whole ThreadX suite and more than every other test in that
configuration put together.

This is the third test in these suites built the same way, after
threadx_thread_priority_change and threadx_thread_wait_abort_and_isr_test: spin
until an interrupt happens to land in a narrow window, with nothing to stop the
spin if it does not. The other two have been given bounds already.

Give this one a wall clock budget too, for the same reason as the last: a tick
is delivered only when the port's timer thread runs, so the tick clock falls
behind real time under load or instrumentation, and instrumentation is exactly
what trace_build turns on.

The check after the loop needs care that the other two did not. It compares
thread_2_counter against thread_2_counter_capture, and the capture is taken
inside the interrupt handler at the moment the window is hit. Leaving that check
in place after a run that never reached the window would compare a live counter
against the zero it was initialised to and report a defect that is not there. So
the check is skipped when the window was not reached, and the run says so.
Reaching the window still exercises it exactly as before.

Verified both ways in trace_build, which is the configuration that was slow: the
window is reached in 8 seconds here and the test passes as it always did, and
with the budget forced to zero the test reports that the window was not reached
and passes without the dependent check firing.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 08:50:15 -04:00
Frédéric DesbiensandClaude Opus 5 4d9ce41845 Gave the wait abort ISR test a budget instead of an open-ended wait (#644)
threadx_thread_wait_abort_and_isr_test waits for an interrupt to land while the
preempt disable flag is set, and waits for it ten times, twenty in the SMP copy,
with no bound on how long that takes. The handler in the same file already says
what can go wrong:

    It is possible for this test to get into a resonance condition in which
    the ISR never occurs while preemption is disabled

and perturbs its own duration to break out of it. That helps but guarantees
nothing, and if the resonance holds, the loop does not end.

It is also, by a wide margin, the most expensive thing in either suite. Run one
test at a time in CI it took between 148 and 726 seconds per configuration:
1936 seconds of the ThreadX suite's 2209, against 273 seconds for the other
ninety five tests together. Nothing else in the suite is within two orders of
magnitude of it.

The budget is in wall clock seconds, not ticks. That distinction turned out to
matter. A tick is delivered only when the port's timer thread gets to run, so
the simulated clock falls behind real time under load or under coverage
instrumentation, and never makes the loss up. A first attempt bounded the wait
at 20000 ticks, nominally 200 seconds, and failed to stop a run that took 726
seconds, because fewer than 20000 ticks had gone by. tx_time_get() cannot bound
elapsed time here; time() can.

A run that falls short says how many windows it reached rather than going quiet,
and the check after the loop is untouched. That check compares semaphore
bookkeeping which holds whatever number of windows were hit, so it still means
exactly what it did before. Hitting the race a few times rather than ten is a
smaller loss than it looks: the value is in reaching the window at all, and the
later hits repeat what the first ones established.

Verified by forcing the budget to 3 seconds, where the test stops after 3.14
seconds of wall clock and reports reaching 0 of 10 windows, with the following
check intact. At 120 seconds both suites pass every configuration run serially.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 08:49:43 -04:00
Frédéric Desbiens b3486f9b46 Bounded the wait in the thread priority change regression tests (#640)
Both copies of this test, the SMP one and the non-SMP one, install an interrupt
handler and then spin until it clears a flag:

    test_isr_dispatch =  test_isr;
    do
    {
        ...
    } while (test_isr_dispatch);

The handler clears that flag only on a narrow window: thread 3 at priority 6,
ready, and not yet at the head of its priority list, which exists only part way
through a priority change. If an interrupt never lands inside that window, the
loop never ends.

That is what has been failing in CI. The SMP suite has been red since 30 June,
and this test times out in three of the last four failing runs, always in a
stack-checking configuration. The evidence that it is a hang rather than slow
work: the test carries no per-test timeout property, so ctest's --timeout 1000
applies, and locally the test finishes in 0.12 seconds with a worst case of 0.29
over thirty runs. Nothing turns that into more than a thousand seconds. It also
survived --repeat until-pass:2, so it hung twice in succession.

The TX_NOT_INTERRUPTABLE path in the same handler already stops after a fixed
amount of work. Only the interruptable path, which is the one the failing
configuration uses, had no protection.

Cap the loop and clear the handler on the way out. When the window is not reached
the test says so and still passes: not reaching it is a gap in what this run
covered, not a fault in the code under test, and failing would report a defect
that does not exist. The counters are left alone so the checks that follow keep
their previous meaning.

The cap is 100000 attempts. An exhausted cap takes about 60 seconds, measured,
and a successful run takes 0.12 seconds, which puts the usual cost around two
hundred attempts and leaves the cap roughly two orders of magnitude clear of it.
That is wide enough not to lose coverage on a slower machine, while replacing a
timeout that says nothing with a message that says what happened.

Not reproducible here, which fits the diagnosis rather than contradicting it: on
sixteen cores the window is hit almost at once. Thirty sequential runs, two
hundred at parallelism thirty-two, sixty pinned to two CPUs, forty pinned to one,
and three full-suite passes at the parallelism CI uses all came back clean. The
defect is the reliance on the window, not any particular machine.

Verified with the window deliberately made unreachable: before this change the
test runs until it is killed, and after it exits in about a minute reporting that
the window was not reached. The full SMP suite passes 110 of 110 at CI's
parallelism.
2026-08-18 17:01:26 -04:00
Frédéric Desbiens 2f6945475b Stopped pthread_self() faulting when the caller is not a pthread (#627)
Nothing prevents an application mixing tx_thread_create() with the POSIX layer,
and a thread created that way has no POSIX control block. posix_thread2tcb()
returns NULL for it, which posix_thread2tid() then read through:

    p_tcb = posix_thread2tcb(thread_ptr);
    thread_ID = p_tcb->pthreadID;

pthread_self() went on to compound it, reading the signal fields of a POSIX_TCB
out of a thread that is only a TX_THREAD:

    if (((POSIX_TCB *) thread_ptr) -> signals.signal_handler)

The first is a null dereference and the second runs off the end of the control
block into whatever the linker put there. Under qemu-system-riscv32 the first one
lands first: mcause=0x5, a load access fault, with mtval=0xb4 for the offset of
pthreadID.

Have posix_thread2tid() report zero for a thread with no control block, which is
what px_pth_join.c already does for the same call, and have pthread_self() skip
the signal check unless the ID says the caller really is a pthread. Zero cannot
collide with a real ID because px_pth_create.c uses the address of the control
block as the ID.

This also covers the case where there is no current thread at all, from an ISR or
before the scheduler starts: tx_thread_identify() returns NULL, and the same
zero comes back instead of a fault.

Add posix_pthread_self_test, which asks both kinds of thread for their ID: a
pthread, which has to report what pthread_create() returned, and a plain ThreadX
thread, which has to report zero. Reverting either half of the fix turns the test
into the load access fault above.

Verified with riscv64-unknown-elf and qemu-system-riscv32: 4 tests across the
default build, 4 of 4 passing.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-16 18:45:16 -04:00
Frédéric Desbiens f184531e7f Added the POSIX compatibility layer to the CMake build, with regression tests (#626)
* Added the POSIX compatibility layer to the CMake build

Nothing in the repository built the POSIX layer. The FreeRTOS layer next to it
has had a target since the CMake build was introduced, so the POSIX one was the
odd one out, and 106 source files went unbuilt by any target, on any
architecture.

Add a posix-threadx target, following the FreeRTOS layer's shape: a static
library, EXCLUDE_FROM_ALL so the default build is unchanged, linking threadx and
publishing its own directory as a PUBLIC include path. The sources live in their
own CMakeLists.txt rather than the top-level file, as common/ does, because there
are 106 of them. The seven posix_*.c files in the same directory are a demo and
standalone signal tests, each with its own entry point, so they stay out of the
library.

The layer does not suit every configuration, and the target is only offered where
it can work:

  - Hosted simulation ports (linux, win32, win64) build against a C library that
    already provides errno.h, pthread.h and the rest. The layer replaces those.
    Its pthread.h even uses _PTHREAD_H, the same include guard as glibc's, so its
    declarations are skipped wholesale and the build fails on missing types.
    Neutralising that guard only exposes the real problem: 69 conflicting
    definitions in a single translation unit, for time_t, struct timespec,
    sigset_t, pthread_t, pthread_mutex_t, sem_t and more. Both the layer and the
    C library implement POSIX, and only one of them can define those names. The
    linux port also emulates threads by calling the C library's pthread_create
    and sem_wait, which the layer exports itself, so linking the two would divert
    the port into the layer that sits on top of it.
  - SMP builds. px_int.h declares _tx_thread_current_ptr as a plain pointer,
    which is a per-core array under SMP, and the layer tracks no current core.

Building the layer for the first time exposed one portability defect worth
fixing rather than working around. tx_posix.h defined ssize_t as INT, with a
comment conceding it should come from <sys/types.h>. That is correct only where
the C library agrees: on AArch64 newlib makes ssize_t 64 bits, and every
translation unit that reached a library header failed to compile. Defer to the
library when it has declared the type, keyed on the _*_DECLARED guards newlib
uses, and do the same for mode_t, which had the same problem waiting. Where no
library declaration exists the previous definitions still apply, so the 32-bit
targets that did build are unaffected.

Verified by building posix-threadx for arm9, arm11, cortex_m0, cortex_m3,
cortex_m4, cortex_m7, cortex_m33, cortex_m55, cortex_m85, cortex_a7, cortex_a9,
cortex_r4, cortex_r5, cortex_a34, cortex_a53 and cortex_a55 with
arm-gnu-toolchain-14.3.rel1, and for risc-v32 and risc-v64 with
riscv64-unknown-elf: 18 of 18, 106 objects each. cortex_a78 has no non-SMP port
and fails to configure with or without this change. Linking the result against
libthreadx.a leaves only tx_application_define, _tx_initialize_low_level, the
optional execution profile hooks, and memset and strlen unresolved, all of which
the application or its C library supplies. The default build still produces
libthreadx.a and no POSIX library.

Compiling is not the same as working, and on the 64-bit targets in that list it
is not enough. The layer carries a message by putting the address of a private
buffer into the queue, and ULONG is 32 bits on every port, so that address only
fits when TX_64_BIT is defined. Without it px_mq_send.c truncates the pointer
and px_mq_receive.c casts the truncated value back, which GCC reports as nothing
worse than a -Wpointer-to-int-cast warning. Defining TX_64_BIT is not a remedy
either: tx_api.h then reaches for the extension pointer macros, which need
tx_thread_extension_ptr in the thread control block, and outside ports_smp and
ports/linux no port declares it. So the target builds everywhere, but the
message queues are only sound on the 32-bit ports. That is pre-existing, it is
not made worse here, and it is left for a change of its own.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Added regression tests for the POSIX compatibility layer

The POSIX layer had no tests. The seven posix_*.c programs shipped beside it are
demos: they print nothing, report no result and end in infinite loops, so they
tell a person watching a debugger something and an automated run nothing.

Add a suite under test/posix, laid out like the FreeRTOS one and driven the same
way, with scripts/build_posix.sh and scripts/test_posix.sh over a run.sh that
takes the same arguments as its RISC-V counterpart.

The tests run on emulated hardware because they have nowhere else to go. The
layer replaces the C library's POSIX headers and exports the same symbols the
linux port calls to emulate threads, so a host build is not available to it. The
RISC-V QEMU harness that the ThreadX suite already uses is, and this suite reuses
its BSP and testcontrol.c rather than growing copies of them.

Three tests to start:

  - posix_mq_basic_test sends a message through a queue and checks the contents
    and priority survive the round trip.
  - posix_mq_send_abort_test covers the leak fixed in #624, by filling a queue,
    blocking a sender on it, aborting the wait and watching the queue's byte
    pool. Reverting the fix makes it fail on the pool check, so it measures what
    it claims to.
  - posix_pthread_basic_test covers pthread creation, a mutex, a semaphore
    handoff, pthread_self and collecting an exit value through pthread_join.

The queue's pool is sized (mq_maxmsg + 1) * (mq_msgsize + 11), which leaves room
for about one message beyond a full queue, so the abort test uses small messages
and a shallow queue. With a larger message the first leaked buffer exhausts the
pool, tx_byte_allocate fails, and the sender disappears into the endless loop in
posix_internal_error() instead of reporting anything. Sizing it this way keeps
the failure legible as a pool measurement rather than a timeout.

riscv32 only, and the reason is the layer rather than the harness. The layer puts
the address of a message buffer into the queue, ULONG is 32 bits on every port,
and a 64-bit address only fits there when TX_64_BIT is defined. Defining it makes
tx_api.h use the extension pointer macros, which need tx_thread_extension_ptr in
the thread control block, and no port outside ports_smp and ports/linux declares
it. Configuring for risc-v64 stops with that explanation rather than building
something that would corrupt a pointer at runtime.

Verified with riscv64-unknown-elf and qemu-system-riscv32: 3 tests across
default_build, disable_notify_callbacks_build, stack_checking_build and
trace_build, 12 of 12 passing.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-16 18:24:48 -04:00
Frédéric Desbiens cfed1d8095 Fixed the kernel object leaks on the static creation error paths (#584)
xQueueCreateStatic() and xTaskCreateStatic() take their storage from the
caller, so neither leaks memory, but both create ThreadX objects and both
return NULL when a later step fails. The caller is left without a handle and
cannot call the matching delete function, so any object already created stays
registered in the kernel, pointing into a caller buffer that the application
is now free to reuse or discard.

Three paths were affected. xQueueCreateStatic() abandoned the read semaphore
when the write semaphore could not be created. xTaskCreateStatic() abandoned
the notification semaphore when the thread could not be created, and abandoned
both the semaphore and the thread when the thread could not be resumed.

Delete what was already created before returning on each of them. The resume
path terminates the thread before deleting it, since a thread created with
TX_DONT_START is suspended rather than terminated, which is the same order the
idle task uses when it reaps a deleted task.

Extend the regression suite to cover all three paths, and add thread resume to
the set of entry points the harness can force to fail. Each static failure case
now uses its own control block, so a future regression on one path cannot carry
damage into the next case and report misleading counts there.

Verified against the layer as it stands on dev, where the three new checks fail
with the objects left behind, and against the fixed layer, where the suite
passes.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-09 08:44:30 -04:00
Frédéric Desbiens f3df5f9dde Added a regression suite for the FreeRTOS compatibility layer (#583)
* Fixed the resource leaks on the xQueueCreate error paths

xQueueCreate() allocated the queue descriptor and its backing memory, then
created two ThreadX semaphores, and returned NULL on either semaphore failure
without releasing anything. Since no handle reached the caller, vQueueDelete()
could not be used to recover, so both allocations were lost. A failure on the
second semaphore additionally abandoned the read semaphore it had already
created, leaving a live ThreadX control block inside freed memory.

Release the backing memory and the descriptor on both paths, and delete the
read semaphore before returning when the write semaphore cannot be created.
This is the teardown order vQueueDelete() already uses, and it matches the
cleanup xTaskCreate() performs on its own error paths.

Verified with a fault injection harness that intercepts the ThreadX byte pool
and semaphore entry points to force tx_semaphore_create() to fail on a chosen
call. On a read semaphore failure the layer previously performed 2 allocations
and 0 releases, and on a write semaphore failure 2 allocations, 0 releases and
0 semaphore deletions. It now performs 2 releases in both cases and deletes
the read semaphore in the second, with the byte pool restored to its prior
state.

Fixes https://github.com/eclipse-threadx/threadx/issues/570

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Added a regression suite for the FreeRTOS compatibility layer

The compatibility layer had no tests in this repository, which is awkward for
its creation functions in particular. Each of them takes one or two byte pool
allocations for its bookkeeping and then creates ThreadX kernel objects, and
each returns NULL when a kernel object cannot be created. The caller is left
without a handle, so it cannot call the matching delete function, and anything
the layer failed to release is gone until the system restarts. A leaking
version and a correct version are indistinguishable from the outside, which is
how the leak in issue 570 went unnoticed.

Add a suite that counts what the layer takes and gives back. A test asks the
harness to fail a chosen kernel creation call, then checks the number of byte
pool allocations, releases, object creations and object deletions performed.
The ThreadX entry points are intercepted with the linker's --wrap so that
tx_freertos.c is compiled exactly as it ships, with no test hooks in it. Note
that tx_api.h maps the public API onto the error checking entry points, so the
_txe_ symbols are the ones wrapped. Coverage is the creation and teardown paths
of queues, tasks, semaphores, mutexes, event groups and timers, including a
regression test for the two paths fixed for issue 570.

The suite follows the layout of the existing ThreadX and SMP suites, is
registered with ctest, and runs in CI through the shared regression template.
It is built 32 bit because the Linux port defines ULONG as unsigned int on
x86_64 while the layer passes pointers through ULONG arguments, so a 64 bit
build truncates them. It is Linux only because --wrap has no MSVC equivalent,
and the CMake configuration says so rather than failing at link time.

Validated by building the suite against the layer as it stands before the
issue 570 fix, where the two expected checks fail with the leaked counts, and
against the fixed layer, where all three tests pass.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-09 08:03:16 -04:00
Frédéric DesbiensandCodex b880ffeada Fixed SMP execution profile total getters (#553)
Updated the SMP execution profile aggregate getters to copy each core's total into the matching output array element. Added a focused regression test for thread, ISR, and idle total getters.

Co-authored-by: Codex <codex@openai.com>
2026-06-22 08:03:46 -04:00
Frédéric DesbiensandCopilot 730b61874b Added copyright headers to files missing them
Applied the standard MIT license header to all project-owned C, header,
assembly, shell, and Python files that were missing a copyright notice.
Third-party, toolchain startup, and auto-generated files were excluded.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-06 21:48:06 +02:00
Frédéric DesbiensandCopilot 223556219+Copilot@users.noreply.github.com 7486de06c8 Refactored, consolidated, and cleaned up RV32/RV64 ports (#536)
risc-v: refactor, consolidate, and fix RV32/RV64 ports

Consolidates the RISC-V 32-bit and 64-bit GNU/Clang port sources, fixes two
pre-existing assembly bugs discovered during testing, and hardens the build
infrastructure for both the regression suite and the CORE-V MCU example.

--- Port consolidation (RV32 GNU + Clang) ---

 - Delete ports/risc-v32/clang/src/ (8 .S files had no Clang-specific
 directives; diverged from GNU only due to missing bug fixes). The Clang
 port CMakeLists.txt now compiles from ../gnu/src/.
 - Change .global -> .weak for _tx_initialize_low_level in gnu/src/ to allow
 BSP-level override without a linker conflict (adopted from Clang port).
 - Create ports/risc-v32/common/tx_port_riscv32_common.h with all definitions
 shared between GNU and Clang ports. Reduce both tx_port.h files to thin
 wrappers.
 - Add a prominent comment in risc-v64/gnu/inc/tx_port.h explaining why
 LONG/ULONG are intentionally 32-bit on RV64 (ThreadX ABI requirement,
 mirrors win64/MSVC LLP64).

--- Shared CMake helper ---

 - Add cmake/threadx_riscv_port.cmake with threadx_add_riscv_port(). All
 three port CMakeLists.txt files are reduced to ~8 lines each. Include path
 is relative to CMAKE_CURRENT_LIST_DIR so the helper works whether ports
 are built standalone or as a subdirectory of the test framework.

--- Shared example-build drivers ---

 - Create canonical driver files under ports/risc-v_common/:
  inc/csr.h                  (uintptr_t-based; portable RV32 + RV64)
  example_build/plic/        (plic.c, plic.h)
  example_build/uart/        (uart_qemu_ns16550.c/h; static inline putc_nolock)
  example_build/trap/        (trap_qemu.c; XLEN-portable mcause constants)
 - Replace per-example copies with symlinks in all qemu_virt and cva6_ariane
 example directories.
 - Fix OS_IS_INTERRUPT typo (was OS_IS_INTERUPT) in shared trap_qemu.c.
 - Gate print_hex() behind TX_RISCV_TRAP_DEBUG.

--- Bug fixes in RV32 assembly ---

tx_thread_schedule.S:

 - Solicited-return FP path: reload t0 from the mepc stack slot before
 csrw mepc, t0. After the FP restore block, t0 held the fcsr value (0 for
 new threads), which caused mepc = 0 and an immediate instruction-address
 fault on the first context switch.
 - Same path: reload t0 from the mstatus stack slot before csrw mstatus, t0
 to avoid writing the stale fcsr value into mstatus.

tx_thread_system_return.S:

 - FP callee-saved registers were saved unconditionally before the mstatus.FS
 check, causing an illegal instruction trap (mcause=0x2) when a thread with
 FS=Off (lazy FPU, thread has never used FP) voluntarily yielded.
 - Apply the same FS guard pattern used in tx_thread_context_save.S: read
 mstatus first, isolate FS[1:0], and skip fsw/fsd if FS == Off.

Both bugs were pre-existing on origin/dev and are unrelated to the
consolidation changes.

--- RV64 64-bit pointer compatibility ---

 - Add TX_TIMER_INTERNAL_EXTENSION, TX_THREAD_CREATE_TIMEOUT_SETUP, and
 TX_THREAD_TIMEOUT_POINTER_SETUP to risc-v64/gnu/inc/tx_port.h to store the
 thread timeout pointer in a VOID
  * extension field rather than truncating it
 into a 32-bit ULONG. Mirrors the win64 port pattern.
 - Define TX_TIMER_EXTENSION_PTR_DEFINED as a portable sentinel.
 - Update threadx_thread_basic_execution_test.c guard from #if defined(_WIN64)
 to #if defined(_WIN64) || defined(TX_TIMER_EXTENSION_PTR_DEFINED).
 - Disable -Wconversion for the RV64 test build: ULONG = unsigned int (32-bit)
 is intentional for ThreadX ABI but triggers spurious warnings when sizeof()
 (8 bytes on RV64) appears in arithmetic with ULONG in common/src/.

--- Regression suite cmake fixes ---

test/tx/cmake/riscv/regression/CMakeLists.txt:

 - Build testcontrol_weak_defaults.c as a separate OBJECT library and include
 it in every test executable via $<TARGET_OBJECTS:>. GNU ld does not extract
 objects from a static archive to satisfy weak symbols, so bundling it in
 test_utility was insufficient for the standalone
 threadx_initialize_kernel_setup_test.

test/tx/cmake/regression/CMakeLists.txt,
test/smp/cmake/regression/CMakeLists.txt:

 - Same fix applied to the Linux and SMP regression builds. The symbols
 abort_all_threads_suspended_on_mutex, suspend_lowest_priority, and
 abort_and_resume_byte_allocating_thread were introduced by the win64 merge
 and left the standalone test unlinkable.

--- CORE-V MCU toolchain and build fixes ---

cmake/riscv64-gcc-rv32imc.cmake:

 - Resolve riscv64-unknown-elf-gcc via PATH so the riscv-collab toolchain in
 /opt/riscv/bin is preferred when it appears first.

ports/risc-v32/gnu/example_build/core_v_mcu/bsp/clz.c (new):

 - The riscv-collab toolchain is built without rv32 multilib, so its libgcc
 does not define __clzsi2 (the helper emitted for __builtin_clz() in fll.c).
 Add a weak __clzsi2 fallback so the build is self-contained with any
 riscv64-unknown-elf toolchain. The weak attribute yields to a
 libgcc-provided strong symbol when the Ubuntu multilib package is used.

core_v_mcu/CMakeLists.txt:

 - Add bsp/clz.c to sources.
 - Reference CMAKE_TOOLCHAIN_FILE via message(STATUS) to suppress the false-
 positive "Manually-specified variables were not used by the project" CMake
 warning and to show the active toolchain at configure time.

--- Housekeeping ---

 - Rename azrtos_test_* -> threadx_test_* (eliminate Azure RTOS branding).
 - Add RV64 QEMU CI test script:
  ports/risc-v64/gnu/example_build/qemu_virt/test/
  threadx_test_tx_gnu_riscv64_qemu.py
 - Normalize entry.s -> entry.S in all 4 example directories.
 - .gitignore: exclude build_m7/ and .codex local artifacts.
 - CI: comment out the riscv regression workflow job and remove it from the
 deploy job's needs list (preserved in-place for easy re-enablement).

--- Verified ---

 - 95/95 RV32 regression tests pass (QEMU virt)
 - 95/95 RV64 regression tests pass (QEMU virt)
 - All 5 Linux build configurations build cleanly (default_build_coverage,
 disable_notify_callbacks_build, stack_checking_build,
 stack_checking_rand_fill_build, trace_build)
 - CORE-V MCU example_build links cleanly with /opt/riscv toolchain

Co-authored-by: Copilot 223556219+Copilot@users.noreply.github.com
2026-05-27 10:30:57 -04:00
Frédéric Desbiens 2c16114a45 Added win64 ports of ThreadX and ThreadX SMP (#529)
Windows x64 port and regression suite

 This PR adds the Windows x64 (Win64) simulation port for both the standalone
 and SMP variants of ThreadX, along with the full CMake build and test
 infrastructure needed to run the regression suite on Windows.


 New ports

 Win64 standalone (ports/win64/vs_2022): self-contained Windows simulation
 port using Win32 threading primitives as virtual cores. Includes CMake
 integration, build/test scripts, and MSVC project files.

 Win64 SMP (ports/win64_smp/vs_2022): multi-core Windows simulation port.
 Supports up to 4 virtual cores backed by Windows host threads.


 Scheduler and timer improvements

 The initial port used coarse polling and synchronous SuspendThread/ResumeThread
 pairs throughout the scheduler hot path. Several rounds of optimization reduced
 the SMP regression suite runtime from ~150 s to ~78 s (-48%), with no
 regressions:

 - Replaced scheduler polling with an event-driven wake path; switched the
   simulated timer to one-shot rearming to eliminate catch-up ticks.
 - Skip SuspendThread when _tx_thread_preempt_disable != 0 (new suspension
   type 3) -- the primary optimization, yielding up to 7.9x speedup on
   preemption-heavy tests.
 - Skip SuspendThread when a thread is spinning on the Win32 critical section
   (suspension type 4), and fix a stale-TLS bug in
   _tx_win32_critical_section_obtain that could stamp mutex_access on the
   wrong virtual core.
 - Added a 2 ms scheduler event timeout (matching the Linux SMP port) to
   prevent stalls on any missed SetEvent.
 - Enabled high-resolution waitable timers (SetWaitableTimerEx) for accurate
   100 Hz tick cadence.
 - Increased TX_WIN32_CONTENTION_PAUSE_COUNT from 64 to 256 to reduce
   SwitchToThread overhead under heavy CS contention.


 Build and test infrastructure

 - Hardened the Windows build wrapper (scripts/build_tx.ps1): invoke Ninja
   directly for Ninja build trees, fix timeout detection, add a default build
   timeout, and limit fallback replay to real timeout cases.
 - Added -Clean support to Windows test scripts to remove stale CTest state
   before each run.
 - Skip Visual Studio DevShell re-entry when the active MSVC environment
   already matches the requested architecture.
 - Fixed scripts/build_tx.sh (Linux) regression source generation: replaced
   brittle exact-string insertion with line-based matching so the interrupt
   dispatcher hook is inserted reliably for both simulator ports.


 Test suite updates

 - Introduced test/tx/regression/threadx_test_port.h with portable macros
   (TX_TEST_POINTER_WORD, TX_TEST_STORE_POINTER) for storing pointers in test
   arrays on 64-bit targets where ULONG remains 32-bit.
 - Adjusted pool-capacity and pointer-storage patterns in regression tests to
   use ALIGN_TYPE-sized slots, making the suite correct on 64-bit hosts.
 - Restored stricter event flag, sleep, and timer expectations now that
   port-level fixes make prior Windows accommodations unnecessary.
 - Tightened SMP watchdog and clean-build timeout defaults.


 Version metadata

 Updated Win32, Win64, and Win64 SMP port version strings to 6.5.1.202602.

 Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
 Co-authored-by: Codex (gpt 5.5) <codex@openai.com>
2026-05-26 17:17:40 -04:00
Akif Ejaz c1e3678797 Add QEMU based CI regression test infra for RV32 and RV64 (#526)
Added a QEMU virt-machine BSP and CTest infrastructure to run the
ThreadX regression suite on both RISC-V 32-bit and 64-bit targets
in CI.

New components:
- BSP (entry, trap, PLIC, CLINT timer, UART, linker script) targeting
  QEMU virt machine for RV32 and RV64
- CMake build system with Ninja, supporting multiple build configs
- CI scripts: install_riscv.sh (toolchain + QEMU), build_tx_riscv.sh,
  test_tx_riscv.sh
- GitHub Actions workflow job for RISC-V regression gating

Port fixes:
- RV32 tx_thread_context_restore.S: set MPIE alongside MPP (0x1800 →
  0x1880) so mret re-enables interrupts
- RV32/RV64 tx_port.h: add TX_REGRESSION_TEST extension macros needed
  by the test harness
- RV32/RV64 example_build scripts: add compile and QEMU launch steps

Regression test portability fixes:
- Block memory tests: increase pool sizes (320 → 340) to accommodate
  larger RISC-V block-header alignment
- Byte memory test: replace hardcoded offsets with BYTE_POOL_OVERHEAD
  macro for portable pool-size computation
- Event flag timeout test: make counter tolerance unconditional,
  removing linux-only guard

Signed-off-by: Akif Ejaz <akif.ejaz@10xengineers.ai>
2026-05-19 11:03:23 -04:00
Frédéric Desbiens 3c4d20285f Corrected typos in two test filenames (#527) 2026-04-29 13:22:15 -04:00
Frédéric Desbiens f81d1c9eb5 Merge branch 'dev' into typo 2026-04-15 11:36:53 -04:00
Frédéric Desbiens c3259a2160 Updated copyright headers and version number constants (#509)
* Updated version number constants

* Removed revision history from all files

* Added Eclipse ThreadX contributors' copyright header
2026-03-05 10:46:30 +01:00
Frédéric Desbiens 1f59529034 Added missing QUEUE_MESSAGE_MAX_SIZE test for SMP. 2025-09-28 22:34:15 +01:00
Haithem Rahmani 87e5110346 Make queue max message size configurable
Summary
-------
This commit fixes the issue #424

Details
--------
- Add a the configuration option TX_QUEUE_MESSAGE_MAX_SIZE in the tx_api.h with default value
  set to TX_ULONG_16 to keep backword compatibility.
- Update the txe_queue_create() to check on TX_QUEUE_MESSAGE_MAX_SIZE rather than TX_ULONG_16
  as max message size.
- Add a new unitary test to cover the new change.
2025-04-02 16:51:36 +01:00
Frédéric Desbiens 898d0ebfde Updated CMake minimal version to 3.13 2025-02-24 15:40:22 -05:00
Yang Hau 9d29a9a8de fix the typos 2024-02-24 00:31:58 +09:00
Yajun Xia e73843f6d4 Added thumb mode support for threadX GNU ports on armv7a platforms. (#333)
* Added thumb mode support for threadX GNU ports on armv7a platforms.

https://msazure.visualstudio.com/One/_workitems/edit/26105175/

* move the swi interrupt to tx_initialize_low_level.S.

* update the test log.
2023-12-28 09:37:39 +08:00
Yajun Xia bc4bd804d5 Fixed the issue of the data/bss section cannot be read from ARM FVP debug tool in cortex-A5 GNU port (#306)
https://msazure.visualstudio.com/One/_workitems/edit/25153813/
2023-09-26 09:51:47 +08:00
Yajun Xia d43cba10b2 Fixed the issue of the data/bss section cannot be read from ARM FVP debug tool in cortex-A9 GNU port. (#303)
https://msazure.visualstudio.com/One/_workitems/edit/25153785/
2023-09-18 16:36:36 +08:00
Yajun Xia a0a0ef9385 Fixed the issue of the data/bss section cannot be read from ARM FVP debug tool in cortex-A8 GNU port. (#302)
https://msazure.visualstudio.com/One/_workitems/edit/25139203/
2023-09-18 10:32:07 +08:00
Yajun Xia 6aeefea8e6 Fixed the issue of the data/bss section cannot be read from ARM FVP d… (#301)
* Fixed the issue of the data/bss section cannot be read from ARM FVP debug tool in cortex-A7 GNU port.

https://msazure.visualstudio.com/One/_workitems/edit/24597276/

* remove untracked files.
2023-09-15 10:46:20 +08:00
yajunxiaMS 7fa087d061 Added thumb mode support under GNU for module manager on Cortex-A7 pl… (#287)
* Added thumb mode support under GNU for module manager on Cortex-A7 platform.

* update code for comment.
2023-07-21 09:26:22 +08:00
Xiuwen CaiandTiejunZhou 6b8ece0ff2 Add random number stack filling option. (#257)
Co-authored-by: TiejunZhou <50469179+TiejunMS@users.noreply.github.com>
2023-05-12 10:13:42 +08:00
TiejunZhou e2a8334f96 Include tx_user.h in cortex_m3/4/7 IAR and AC5 port (#255)
* Include tx_user.h in ARMv7-M IAR port

* Include tx_user.h in ARMv7-M AC5 port

* Include tx_user.h in cortex_m3/4/7 IAR and AC5 port
2023-04-24 09:33:00 +08:00
TiejunZhou 7a3bb8311b Release scripts to validate ThreadX port (#254) 2023-04-23 10:58:21 +08:00
TiejunZhou 4c4547d5d5 Fix path to test reports in pipeline (#247)
* Fix path to test reports in pipeline

* Fix test case when CPU starves, the thread 2 can run 14 ronuds.
2023-04-17 09:40:59 +08:00
TiejunZhou 0d308c7ae6 Fix random failure in test case threadx_event_flag_suspension_timeout_test.c (#246)
Depending on the starting time, thread 1 can run either 32 or 33 rounds.
2023-04-14 14:55:04 +08:00
Tiejun Zhou 5f430f22e2 Add Azure DevOps pipelines for ThreadX test 2023-04-12 09:40:17 +00:00
Tiejun Zhou ebeb02b958 Release ThreadX regression system 2023-04-04 09:40:54 +00:00