3 Commits
Author SHA1 Message Date
Frédéric Desbiens 9a03838381 Stopped the test runners from reporting an incomplete build as test failures (#709)
The cmake test runners drove Ninja with its default keep-going of 1, so the
first failing target ended the build. Every target scheduled after it was
simply absent, and ctest reports a missing binary as a failing test. A
single link error therefore produced a failure count that moved with build
scheduling order rather than with the code.

Measured on the RISC-V64 regression suite, where two targets genuinely
cannot link:

    before   79 of 95 test binaries built
    after    93 of 95 test binaries built

Fourteen perfectly good binaries were being skipped and counted as
failures. Passing -k 0 lets Ninja finish everything it can; the build still
exits non-zero when a target fails.

Three related problems in the same paths are fixed with it.

A failing configuration used to abort the loop over configurations, so
under set -e the ones after it went unbuilt or untested. The build loops
and the serial test loops now accumulate status and return it at the end,
which is what the parallel test branch already did with wait, and what
cmake_bootstrap.sh already documented for ctest.

Capturing that status removes the set -e protection inside the functions,
so two latent faults become reachable and are closed here. A failed pushd
would have let ctest run in the source tree, where it finds no tests and
reports success; the pushd is now guarded. And ctest's status was
discarded by the popd that follows it, so a configuration with failing
tests returned 0 and was reported as a pass; the status is now carried
past the popd and the summary steps.

Verified on the RISC-V32 suite, which has two genuine failures in each of
its five configurations. Both the serial and the parallel branch now test
all five and exit 8, where the serial branch previously stopped after the
first configuration.

The tx and smp runners are symlinks to scripts/cmake_bootstrap.sh, so they
are covered by the one change there.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-09 10:56:28 -04:00
Frédéric Desbiens 2f6945475b Stopped pthread_self() faulting when the caller is not a pthread (#627)
Nothing prevents an application mixing tx_thread_create() with the POSIX layer,
and a thread created that way has no POSIX control block. posix_thread2tcb()
returns NULL for it, which posix_thread2tid() then read through:

    p_tcb = posix_thread2tcb(thread_ptr);
    thread_ID = p_tcb->pthreadID;

pthread_self() went on to compound it, reading the signal fields of a POSIX_TCB
out of a thread that is only a TX_THREAD:

    if (((POSIX_TCB *) thread_ptr) -> signals.signal_handler)

The first is a null dereference and the second runs off the end of the control
block into whatever the linker put there. Under qemu-system-riscv32 the first one
lands first: mcause=0x5, a load access fault, with mtval=0xb4 for the offset of
pthreadID.

Have posix_thread2tid() report zero for a thread with no control block, which is
what px_pth_join.c already does for the same call, and have pthread_self() skip
the signal check unless the ID says the caller really is a pthread. Zero cannot
collide with a real ID because px_pth_create.c uses the address of the control
block as the ID.

This also covers the case where there is no current thread at all, from an ISR or
before the scheduler starts: tx_thread_identify() returns NULL, and the same
zero comes back instead of a fault.

Add posix_pthread_self_test, which asks both kinds of thread for their ID: a
pthread, which has to report what pthread_create() returned, and a plain ThreadX
thread, which has to report zero. Reverting either half of the fix turns the test
into the load access fault above.

Verified with riscv64-unknown-elf and qemu-system-riscv32: 4 tests across the
default build, 4 of 4 passing.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-16 18:45:16 -04:00
Frédéric Desbiens f184531e7f Added the POSIX compatibility layer to the CMake build, with regression tests (#626)
* Added the POSIX compatibility layer to the CMake build

Nothing in the repository built the POSIX layer. The FreeRTOS layer next to it
has had a target since the CMake build was introduced, so the POSIX one was the
odd one out, and 106 source files went unbuilt by any target, on any
architecture.

Add a posix-threadx target, following the FreeRTOS layer's shape: a static
library, EXCLUDE_FROM_ALL so the default build is unchanged, linking threadx and
publishing its own directory as a PUBLIC include path. The sources live in their
own CMakeLists.txt rather than the top-level file, as common/ does, because there
are 106 of them. The seven posix_*.c files in the same directory are a demo and
standalone signal tests, each with its own entry point, so they stay out of the
library.

The layer does not suit every configuration, and the target is only offered where
it can work:

  - Hosted simulation ports (linux, win32, win64) build against a C library that
    already provides errno.h, pthread.h and the rest. The layer replaces those.
    Its pthread.h even uses _PTHREAD_H, the same include guard as glibc's, so its
    declarations are skipped wholesale and the build fails on missing types.
    Neutralising that guard only exposes the real problem: 69 conflicting
    definitions in a single translation unit, for time_t, struct timespec,
    sigset_t, pthread_t, pthread_mutex_t, sem_t and more. Both the layer and the
    C library implement POSIX, and only one of them can define those names. The
    linux port also emulates threads by calling the C library's pthread_create
    and sem_wait, which the layer exports itself, so linking the two would divert
    the port into the layer that sits on top of it.
  - SMP builds. px_int.h declares _tx_thread_current_ptr as a plain pointer,
    which is a per-core array under SMP, and the layer tracks no current core.

Building the layer for the first time exposed one portability defect worth
fixing rather than working around. tx_posix.h defined ssize_t as INT, with a
comment conceding it should come from <sys/types.h>. That is correct only where
the C library agrees: on AArch64 newlib makes ssize_t 64 bits, and every
translation unit that reached a library header failed to compile. Defer to the
library when it has declared the type, keyed on the _*_DECLARED guards newlib
uses, and do the same for mode_t, which had the same problem waiting. Where no
library declaration exists the previous definitions still apply, so the 32-bit
targets that did build are unaffected.

Verified by building posix-threadx for arm9, arm11, cortex_m0, cortex_m3,
cortex_m4, cortex_m7, cortex_m33, cortex_m55, cortex_m85, cortex_a7, cortex_a9,
cortex_r4, cortex_r5, cortex_a34, cortex_a53 and cortex_a55 with
arm-gnu-toolchain-14.3.rel1, and for risc-v32 and risc-v64 with
riscv64-unknown-elf: 18 of 18, 106 objects each. cortex_a78 has no non-SMP port
and fails to configure with or without this change. Linking the result against
libthreadx.a leaves only tx_application_define, _tx_initialize_low_level, the
optional execution profile hooks, and memset and strlen unresolved, all of which
the application or its C library supplies. The default build still produces
libthreadx.a and no POSIX library.

Compiling is not the same as working, and on the 64-bit targets in that list it
is not enough. The layer carries a message by putting the address of a private
buffer into the queue, and ULONG is 32 bits on every port, so that address only
fits when TX_64_BIT is defined. Without it px_mq_send.c truncates the pointer
and px_mq_receive.c casts the truncated value back, which GCC reports as nothing
worse than a -Wpointer-to-int-cast warning. Defining TX_64_BIT is not a remedy
either: tx_api.h then reaches for the extension pointer macros, which need
tx_thread_extension_ptr in the thread control block, and outside ports_smp and
ports/linux no port declares it. So the target builds everywhere, but the
message queues are only sound on the 32-bit ports. That is pre-existing, it is
not made worse here, and it is left for a change of its own.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Added regression tests for the POSIX compatibility layer

The POSIX layer had no tests. The seven posix_*.c programs shipped beside it are
demos: they print nothing, report no result and end in infinite loops, so they
tell a person watching a debugger something and an automated run nothing.

Add a suite under test/posix, laid out like the FreeRTOS one and driven the same
way, with scripts/build_posix.sh and scripts/test_posix.sh over a run.sh that
takes the same arguments as its RISC-V counterpart.

The tests run on emulated hardware because they have nowhere else to go. The
layer replaces the C library's POSIX headers and exports the same symbols the
linux port calls to emulate threads, so a host build is not available to it. The
RISC-V QEMU harness that the ThreadX suite already uses is, and this suite reuses
its BSP and testcontrol.c rather than growing copies of them.

Three tests to start:

  - posix_mq_basic_test sends a message through a queue and checks the contents
    and priority survive the round trip.
  - posix_mq_send_abort_test covers the leak fixed in #624, by filling a queue,
    blocking a sender on it, aborting the wait and watching the queue's byte
    pool. Reverting the fix makes it fail on the pool check, so it measures what
    it claims to.
  - posix_pthread_basic_test covers pthread creation, a mutex, a semaphore
    handoff, pthread_self and collecting an exit value through pthread_join.

The queue's pool is sized (mq_maxmsg + 1) * (mq_msgsize + 11), which leaves room
for about one message beyond a full queue, so the abort test uses small messages
and a shallow queue. With a larger message the first leaked buffer exhausts the
pool, tx_byte_allocate fails, and the sender disappears into the endless loop in
posix_internal_error() instead of reporting anything. Sizing it this way keeps
the failure legible as a pool measurement rather than a timeout.

riscv32 only, and the reason is the layer rather than the harness. The layer puts
the address of a message buffer into the queue, ULONG is 32 bits on every port,
and a 64-bit address only fits there when TX_64_BIT is defined. Defining it makes
tx_api.h use the extension pointer macros, which need tx_thread_extension_ptr in
the thread control block, and no port outside ports_smp and ports/linux declares
it. Configuring for risc-v64 stops with that explanation rather than building
something that would corrupt a pointer at runtime.

Verified with riscv64-unknown-elf and qemu-system-riscv32: 3 tests across
default_build, disable_notify_callbacks_build, stack_checking_build and
trace_build, 12 of 12 passing.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-16 18:24:48 -04:00