mirror of
https://github.com/eclipse-threadx/threadx.git
synced 2026-10-06 06:59:08 +08:00
ee4c2fb6c7909521cc6ea42d5d7c22f103707c52
64
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ee4c2fb6c7 |
Merge commit from fork
A memory-protected module's object lookup arrives at the Module Manager's dispatch layer, which decides how much of the module's name buffer the privileged comparison may touch. It checked one byte. _txm_module_manager_object_name_compare reads a character from each name at the top of every iteration and tests the remaining search length at the bottom, so a declared length N permits reads at offsets 0 through N: N+1 bytes, the extra one being the terminator a length excludes. The extended service receives that length in its extra parameter array, validates the array correctly, and then never uses the length to bound the buffer it describes. The deprecated service carries no length at all; its manager wrapper calls the same implementation with the largest value a UINT can hold, so it removes the bound rather than lacking one. The comparison stops at the first differing character, so a walk past the end of the module's memory looks as though it needs the bytes there to match a name the module chose. It does not. A create service stores the name pointer a module supplies in the control block rather than copying the string, so a module can register an object whose name is the very buffer it then searches for. Both sides of the comparison are then the same address, every character matches by construction, and the walk continues until it meets a terminator in memory the module does not own and cannot see. That walk runs in privileged mode with _tx_thread_preempt_disable raised, and none of the 27 memory fault handlers lowers it again, so a fault on it costs more than the requesting thread. The fix validates the range the comparison may reach. The extended dispatcher now checks the name over name_length + 1 bytes, refusing a length whose range cannot be expressed, and the extra parameter array is checked first because the length comes out of it. The deprecated dispatcher refuses a memory-protected module outright, because no range can be derived from a pointer alone; a module without protection is unchanged, as it is for every other check in this layer. _txm_module_manager_object_name_compare is left as it is. Reading N+1 bytes for a declared length of N is the documented contract, and the defect is that the contract was never checked. The regression test drives both lookup dispatchers and measures two extents rather than sampling either side of a boundary, because the finding is the relationship between them. It anchors the name buffer to the end of one of the module's regions and walks it backwards to find the smallest room the dispatcher accepts, which is the extent validated; and it fills memory with a filler character, plants the only terminator at a chosen depth, and asks for an object named with exactly the bytes up to it, so that the deepest depth the lookup can be made to return from is the extent read. Against dev's copy of the two headers the test exits 1 with 408 failed expectations of 1130: a 31-character name with one byte of room is accepted and read 31 bytes past the region, and the aliased walk reaches every depth offered. With the fix it exits 0 with 1490 expectations and none failed. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
a967d005f6 |
Merge commit from fork
_txm_module_manager_tx_block_pool_create_dispatch proved that two ALIGN_TYPE words of the extra-parameter array a module supplies lay inside the module's data, and then used indices 0 through 3. The two words it had not proved were read twice each in privileged context: once by the dispatcher's own buffer check, which takes index 1 as the pool start and index 2 as the length to range-check it over, and once when the argument list for _txe_block_pool_create was built. So the out-of-bounds read was performed by the validation code itself, and the size the caller's pool start was then validated against was a word the manager had not proved the module owned. A module reaches this by calling its kernel dispatcher directly with an array that ends two words before the boundary of its data or of a shared region, which the module library's own four-word array never does. The read window is eight bytes on every module port and there is nothing in it the module can lengthen. The extent becomes sizeof(ALIGN_TYPE[4]), which is what the two siblings of the same shape have always had: queue create validates four words for four, and byte-pool create three for three. Both checks on this path are data-only, with no fallback to the portable code-region check, so the four Cortex-A35 and Cortex-A35 SMP module port combinations whose data check is the constant (TX_SUCCESS) refuse these create requests from a memory-protected module before and after this change alike. The regression test drives all three create dispatchers and measures the extent each one validates rather than sampling either side of it: it anchors the array at the end of each region a module owns and walks it backwards until the dispatcher accepts, and the smallest room accepted is the extent. It measures the highest index each dispatcher uses through what the stubbed service records, so the invariant the three are checked against is measured on both sides. It also measures what a module can learn from the answers, since the buffer check is a comparison against a threshold the module chooses and the dispatcher's refusal is distinguishable from every status these services return. Compiled against the previous dispatch header the test reports the extent as two words and the service reached carrying a word planted past the end of the region; against this one it reports four. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
35d4137126 |
Merge commit from fork
A module calls txm_module_object_allocate and chooses object_size. The manager adds sizeof(TXM_MODULE_ALLOCATED_OBJECT) to it and asked the object pool for the result without checking the addition, so a size near the top of a ULONG wrapped: object_size 0xFFFFFFF8 asked for eight bytes, the pool served them, and the manager then wrote its sixteen-byte header -- owner, list links and size -- into that eight-byte block and linked it into the module's allocation list. Eight bytes of the next block's contents or header are gone by the time the call returns, and the shared object pool that every module allocates from is the thing that was corrupted. The addition is now made with the repository's own overflow-checked helper, which was already used on both load paths and simply never reached this one, and it is made before the protection mutex is taken so a refused request leaves the pool, the allocation list and its count exactly as they were. Sizes that survive the addition need no further bound here: _txe_byte_allocate refuses a request larger than the pool, so the alignment round-up in _tx_byte_allocate is never reached with a value that could wrap in its turn. Regression coverage runs on the host by modelling the byte pool behind the manager. The model applies the two bounds _txe_byte_allocate applies and hands out a block of precisely the requested length with a guard band immediately after it, so an undersized allocation is caught as the out-of-bounds write it is rather than inferred from arithmetic. Thirty-two of its expectations fail against the previous implementation, twenty reporting state a refused request must not have touched and two reporting the eight and twelve bytes written past the end of the block. It reaches 100% line, branch and call coverage of the changed function; the lines it leaves uncovered are its own failure reporting, plus the modelled pool's zero-size refusal, which only an unfixed build reaches. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
ff4be30656 |
Covered the control block ID cleared when module memory is given back (#779)
The control block ID that _txm_module_manager_object_deallocate clears on its way out is read by the size-based object check, which is the one kept for dispatchers outside this repository. Nothing exercised that path. The expectations covering the clearing all went through object authentication, which consults the kernel's created list and would refuse a recycled address whatever its ID said, so they would have passed with the clearing removed. A section drives the size-based check directly. A module plants the value of an ID into memory it owns, gives the allocation back, and is handed the same address again; the check accepts the planted ID before the deallocation and refuses it afterwards, which is the difference the clearing makes. The authentication test holds 125 expectations and passes. Removing the ID clearing fails it. The full suite passes 107/107 in default_build_coverage with no warnings. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
2930618cfe |
Merge commit from fork
* Refused to give back the memory of a live kernel object A module allocates the control blocks of its kernel objects from the Module Manager's object pool and can ask for that memory back by address. The manager released it whatever was in it, including a control block the kernel was still using. Deallocation is not deletion: the object stays on the created list for its type, a thread stays wherever it was on the ready or suspension lists and stays schedulable, and an active timer stays on the timer list. Nothing on those paths consults a control block ID, so nothing about the memory having been freed stops the kernel from using it -- a created list walk reads a name pointer out of it and follows that pointer, a create or delete of another object of the same type writes through the created links in it, the scheduler switches the stack pointer to the word at offset 8 of it and pops a saved processor state, and timer expiration calls the function pointer in it. Meanwhile the byte pool is free to hand those same bytes to the next allocation, so what the kernel goes on reading as a control block becomes whatever the next owner of the memory puts there, through ordinary create services. The manager now refuses to release memory that holds an object which is still created, and returns TX_DELETE_ERROR without touching the allocation, the object or the memory. Deleting the object first is what makes its memory releasable, which is the sequence the delete dispatchers already follow and the one module authors are told to follow. The question is answered from the kernel's created lists, not from the control block. A control block ID is not evidence that an object is there: an allocation that was never created can be carrying the value of an ID, and refusing on that would strand memory a module is entitled to have back. Deallocation is also given an address and nothing else, so unlike a typed service it cannot be told which list to search, and each of the eight lists is searched in turn. The search is not narrowed by the size of the allocation, because that would be sound only if every object had been created through a size-checked path, and a module running without memory protection creates objects through no such path -- a queue at the start of a thread-sized allocation is a case the tests here cover. Each list is searched in its own interrupts-disabled window, bounded by the count the kernel keeps beside it, so the longest window is the length of one type's list and a list whose links have been damaged cannot make the search run on. The whole search is one pass over the objects the system has created, paid once per object deallocation. Storage that is not an object is released exactly as before, which is what keeps cleanup after a create that failed or was abandoned working, and what makes the release each delete dispatcher performs after a successful delete go through. The address the request arrives with is now checked before the manager's private header in front of it is read, rather than partly after. The size of the allocation comes from that header, so there is nothing to validate a size against until the header has been read, and the previous order established only where the header started: a header that began inside the pool and ended past it had its size word read from outside the pool. That check has moved out of the dispatcher into _txm_module_manager_param_check_object_for_deallocation, alongside the other parameter checks the dispatch table uses and where a test can reach it, and it now also refuses a size that would carry the end of the allocation past the top of the address space instead of wrapping it, since a wrapped end compares as though the allocation were inside the pool. The 301 expectations in the new test drive all eight object types through allocate, create, a refused deallocation, delete, and a deallocation that succeeds, and assert after the refusal that nothing reached the pool, that the allocation is still on the module's list at the head of it, that the control block still carries its ID and that the object is still live. They cover storage that was never created, storage carrying nothing but a plausible ID for each of the eight types, an object at the start of an oversized allocation, every aligned interior offset of a live object with that object's own ID planted at it, an application-owned object outside the pool, another module's allocations both live and raw, releasing the head of a list of several and releasing the same address twice, a type that has no created list, the bounds on the search against a list longer than its count and against a count larger than its list for every type, the boundary addresses at both ends of the pool, a crafted size, and an object pool that was never created. Removing the refusal fails 43 of them; removing the size wrap guard fails one. Line and branch coverage of the three new functions and of the changed _txm_module_manager_object_deallocate is 100%, with one exception that is test scaffolding rather than product code: the host shim's stand-in for TX_RESTORE has an underflow guard, and the test asserts that branch is never taken. All 99 tests pass in each of the five configurations the tree builds with GCC 14. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> * Removed the object deallocation bounds check superseded by the search _txm_module_manager_param_check_object_for_deallocation() bounded the private header in front of a caller's address before the deallocator read it. The deallocator no longer reads that header: it finds the allocation by searching the module's own allocation list, which never dereferences the address, so the bounds check now guards a read that does not happen. The function, its prototype, its macro and its one call site in the txm_module_object_deallocate dispatcher are removed, along with the twelve expectations that covered it and three declarations left unused by their removal. The search proves more than the check did: the bounds test established only that the header lay inside the object pool, while the search establishes that the address is the exact start of one of this module's allocations. The live object deallocation test holds 289 expectations and passes. Reverting the live-object guard still fails 51 of them, the same 51 as before the deletion, so nothing the removed expectations covered was load-bearing. The full suite passes 108/108 in default_build_coverage, with no warnings. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
3276b0efd5 |
Merge commit from fork
A module reaches a kernel object's delete service through the Module Manager's dispatch layer, which asked one question about the pointer: does the object lie outside the module's own code and data. That is the policy for using an object, and using an object a module does not own is supported and intended -- txm_module_object_pointer_get_extended searches the system's created objects by name and hands back objects the application created and objects other modules created, so that a module can send to a shared queue, take a shared semaphore or get a shared mutex. Destroying one of those is not sharing it. A module could name any object it had a pointer to and have the privileged dispatcher delete it: a host service lost its queue, waiters were resumed with TX_DELETED, an owned mutex's priority inheritance was unwound, and an active timer stopped. Nothing in the manager intended that. The deallocation the delete path performs afterwards has always required the memory to belong to the calling module, so a delete of anything else could only ever end in TX_PTR_ERROR -- the ownership requirement was already there and was enforced one step too late, after the kernel object had been irreversibly destroyed and its waiters woken. An error returned at that point describes a cleanup failure and restores nothing. The eight delete dispatchers now ask the question before the delete rather than after it. A memory-protected module may delete an object only if the object is one it allocated from the manager's object pool, at the exact address the manager returned to it, for an allocation made for that type's control block size -- the same conditions the create dispatchers already apply, since deletion is the inverse of creation. Anything else returns TXM_MODULE_INVALID_MEMORY, which is what every other parameter rejection in the dispatch table returns, so no new value enters the module ABI and a module cannot use a delete request to tell "not yours" apart from "not a usable address". The object stays created, nothing waiting on it is resumed, and no created count moves, because the kernel was never asked. The ownership question is deliberately separate from whether an object is there at all and of the expected type. It establishes nothing about liveness or type, and it is composed with the checks that do rather than replacing them. _txm_module_manager_object_deallocate reached the manager's private header by subtracting from whatever address the caller supplied, and read it before deciding whether it was a header. All eight delete dispatchers call that function directly, so an object the module does not own arrived there as a matter of course rather than exceptionally: for an object the application allocated statically, or for any address outside the object pool, the words in front of it were unrelated memory read in privileged mode -- and acted on, since an address whose preceding words happened to name the calling module was unlinked from that module's allocation list and handed to the byte pool. It now finds the allocation by searching the module's own allocation list. The caller's address is compared and never dereferenced, so an address that names none of this module's allocations is refused without a privileged read of anything in front of it, and the search establishes what the header read could not: that the address is the exact start of an allocation rather than somewhere inside one. The search is bounded by the count the manager keeps beside the list and runs with interrupts disabled, so a damaged list cannot make it run on and it cannot observe the list being changed under it. The fix is in the deallocation itself rather than at a call site, so it covers all eight delete dispatchers, the deallocation request a module can make directly, protected and unprotected modules alike. txm_module_manager_stop needs no exemption and was not given one. It deletes the objects a module created by calling the internal _tx_*_delete services directly, and identifies them with _txm_module_manager_created_object_check rather than through a module request, so no caller-facing check stands in its way. The 209 expectations in the new test drive all eight object types through allocation, an ownership check by the owning module and by another, a check against a larger and a smaller expected size, and a deallocation that succeeds. They cover an object the application owns, deliberately preceded by a header naming the requesting module so that reading it can be seen to have happened; another module's object, checked and refused from both sides; every aligned interior offset of an allocation and the addresses of its header, its end, the pool's ends and one past the pool; a null pointer and an address with no room for a header in front of it; a request from no module at all; memory that has changed hands, where a stale pointer names an address that now belongs to another module; the bounds on the search, against a count smaller than the list and against a count larger than a list whose links are broken; releasing the head, the middle and the last of a list of three; and an object pool that was never created. Every call asserts that the interrupt lock came back balanced, and that the search ran with interrupts disabled exactly one deep. Restoring the previous deallocation fails four of them and then segmentation faults, on the null-pointer case, where the header in front of address zero is read; removing the ownership check fails 74. Line and branch coverage of the two new functions and of the changed _txm_module_manager_object_deallocate is 100%, with two branch outcomes excepted that are test scaffolding rather than product code: the host shim's stand-ins for TX_DISABLE and TX_RESTORE carry a maximum-depth test and an underflow guard, and the real ports' primitives are inline assembly with no branch at all. All 99 tests pass in each of the five configurations the tree builds with GCC 14. Nothing in the tree compiles the dispatch header, so the eight changed dispatchers were cross-compiled by hand: the two changed C files and a translation unit that includes the header build at -Werror with arm-none-eabi-gcc 13.2.1 for every GNU 32-bit module port -- Cortex-A7, M0+, M23, M3, M33, M4 and M7 -- and produce the same -Wall -Wextra warning counts as before the change, 117 on Cortex-A7 and 116 on each of the others. Cortex-R4 and RXv2 have no GNU module port, and the two AArch64 module ports have no toolchain available here; the new arithmetic is a comparison of two addresses of the same type, so a wider ALIGN_TYPE changes nothing about it. MISRA: all if bodies are braced, the address comparison goes through ALIGN_TYPE, no pointer the caller supplied is dereferenced, and no goto appears. No deviation is required. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
7495b02374 |
Merge commit from fork
A memory-protected module names the kernel objects it wants operated on by address, and the Module Manager decided whether a privileged service could dereference that address by asking only whether it fell outside the module. The object pool is outside every module, so that test was satisfied by an address shifted into the interior of one of the module's own privileged allocations, which denotes no object at all. The bytes such an address presents as a control block are bytes the module put there through ordinary create and set services, so the control block ID at the front of them could be made to read as any type the module chose, and the _txe_ layer's ID test then agreed. A module could therefore have the kernel read and write fields of an object that does not exist, at an address it picked; the reported chain reaches a privileged memset across an attacker-chosen range that way. Authentication now answers the question the location test could not: is this the exact address of a live kernel object of the type this service expects. It is answered from the kernel's created list for the type, which the create and delete services maintain, so membership establishes at once that the address is an object start rather than an address inside an object, that the object is of this type, and that it has not been deleted. That list is also what makes an object the application created and shared with a module authenticable at all, since the manager allocated no such object and has no record of its own to consult, so sharing keeps working and needs no new registration API. Nothing a module can influence is used to establish a type. An earlier form of this fix accepted an address that was the exact start of one of the calling module's own allocations of the right size and then took the type from the control block ID, which is cheaper -- a module's allocation list is much shorter than the system's created list for a type. The tests here refused it: a byte pool, a mutex and a timer are all the same size on a 32-bit target, so an allocation created as one of them and presented as another passed both the address and the size test, leaving the ID as the only thing between the module and a type confusion. The ID is still checked, after the created list has settled the question, because it is the test the _txe_ services make and it is compiled away with them under TX_DISABLE_ERROR_CHECKING. An address a module named is not read at all until a kernel record says there is an object there. All 58 object-using dispatchers now pass the object type rather than a control block size, since a size cannot establish a type. Both scans are bounded by the counts the kernel and the manager maintain beside their lists, so the cost of one check is bounded and a list whose links have been damaged cannot make a scan run on, and both run with interrupts disabled, which is the protection those lists are maintained under and which adds no blocking point or priority inversion to a kernel request. The cost is a walk of the created list for the type, paid only by memory-protected modules; the size-based check is kept for dispatchers outside this repository and hardened to refuse object pool interiors, which is the part of the attack it can see without a type. Two further changes close paths the authentication alone would leave open. Thread reset now refuses a thread that carries no module instance: such a thread is authentic, can be found by name and satisfies reset's own state test, and reset reads the shell entry function out of that instance, so a null one was followed in privileged mode. Object deallocation now clears the control block ID as it gives memory back, so an object freed without being deleted first stops being vouched for at the moment the manager stops owning its memory, rather than returning to the pool still carrying a valid ID for the next allocation placed there to present. The 120 expectations in the new test offer every aligned interior offset of a legitimate object as a foreign type, with that type's own ID planted at the offset, and assert that none is accepted; they cover all eight object types against each other, uncreated allocations, deleted objects, freed addresses that have been handed out again, objects the application owns and memory that merely carries a plausible ID, objects another module allocated, and the bounds on both scans. Reverting the object checks to the permissive policy fails 89 of them; removing the thread reset guard makes the test die with SIGSEGV inside the reset path; removing the ID clearing in deallocation fails 2. All 99 tests pass in each of the five configurations the tree builds with GCC 14. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
b37cd4a81a |
Fixed RISC-V regression portability failures (#773)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
RISC-V regression builds ran only at -O0, leaving a timer callback counter that spins forever at -O2. The trace regression was excluded because it referenced a port-specific interrupt-save variable, and its adjacent pools could start misaligned on RV64. Made the timer counter volatile, gave the trace test aligned pool storage and a portable saved-interrupt value, and enabled it on RISC-V. Added an -O2 QEMU configuration to keep the optimized failure covered. CMake/Ninja/QEMU: RV64 default, optimized and trace suites passed 97/97 each; RV32 passed 96/96 each. Two ISR event tests passed 30 repeats each at -O2. The Linux/GCC 14 trace test passed. The reported timing resonance did not recur. Assisted-by: Codex (gpt-6-sol) <noreply@openai.com> |
||
|
|
ad558a7b1f |
Fixed the zero trace time stamps in the Linux ports' MISRA builds (#749)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
Both Linux ports define TX_TRACE_TIME_SOURCE as _tx_misra_time_stamp_get() when TX_MISRA_ENABLE is set, and neither implements that function, so both inherit the generic `return(0);` from tx_misra.c. Every trace event is stamped zero. The buffer carries no timing, and the kernel's own check for an entry having been overwritten -- time_stamp against the entry's own stamp, in the block and byte allocates and in the system suspend and resume -- compares zero with zero, so it never fires and a service patches whatever now occupies the slot. Both ports now read in MISRA builds the clock they already read otherwise, _tx_linux_time_stamp.tv_nsec, which TX_TRACE_PORT_EXTENSION refreshes on every recorded event in both forms of the insert. The non-SMP port's non-MISRA macro carried a trailing semicolon, which made it a statement and is why the MISRA insert -- which takes the time source as a function argument -- could not use it; that is dropped and the two branches become one definition. Both headers keep the _tx_misra_time_stamp_get declaration, because tx_misra.c still defines it and is compiled for these ports. The MISRA insert evaluates its time source before the callee refreshes the clock, so each entry carries the reading taken at the previous recorded event. Stamps are real, distinct and ordered, which is what the overwrite check needs. The trace entry update test gains an assertion that the buffer holds an entry the port actually stamped. It fails on dev with ERROR #13 under misra_trace_build and passes with this change. Suites green: 7/7 ThreadX configurations, 5/5 SMP. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
e9b5aec65e |
Added MISRA build configurations to the ThreadX regression matrix (#747)
TX_MISRA_ENABLE routes the kernel's pointer conversions, its memset and its stack check through the shim in common/src/tx_misra.c. No build configuration defined it, so that file was never compiled or tested, and the source the MISRA analysis runs against was not the source the suite exercises. Adding event tracing alongside it reaches a further region of the shim that neither macro alone compiles. Both defects found in this area were invisible for the same reason: #741 broke under TX_MISRA_ENABLE and #745 under both macros together, and neither combination was built anywhere. misra_build and misra_trace_build fill that gap. Two things had to give way for them. TX_POINTER_TO_ALIGN_TYPE_CONVERT exists only when TX_MISRA_ENABLE is absent, so threadx_test_port.h writes the conversion out rather than taking it from the API; it is test scaffolding storing a pointer in a word wide enough to hold one, not kernel source. And thread_transition and module_manager compile hand-picked common/src sources directly instead of linking the library, which is what lets them reach feature macros the library-per-configuration model cannot; under TX_MISRA_ENABLE those sources call into the shim, and pulling the shim in pulls its own callees after it. Neither directory exists to test the shim, so the MISRA configurations skip them. 100/100 tests pass in each of the two new configurations, against 105/105 in the existing five - the five not run are the four thread transition tests and the module manager test, exactly the two directories skipped. Build is 4 to 5 seconds and the suites 13 and 15 seconds, against 4 seconds and 9 to 17 for the configurations already there, so the tx job grows by roughly forty seconds. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
24d1f278c2 |
Decorrelated the wait abort ISR test's interrupt source from the thread it samples (#743)
threadx_thread_wait_abort_and_isr_test waits for the periodic timer interrupt to land while the preempt disable flag is set. The same interrupt also puts the semaphore that wakes the thread being sampled, so the two run in lockstep: the tick wakes thread 0, thread 0 does a fixed amount of work and suspends, and the next tick arrives a fixed interval later with thread 0 in the same place every time. That is the resonance the handler's own comment describes, and perturbing the handler's duration only shifts the phase rather than breaking the correlation. The window itself is a few instructions wide, so a sampler locked to the thread's own cycle can miss it indefinitely, which is what #644 and #649 measured and worked around. The simulator's timer thread waits on _tx_linux_timer_semaphore with a one-tick deadline and delivers an interrupt early when the semaphore is posted, which the port already relies on in _tx_thread_schedule. The test now runs a plain POSIX thread that posts it, injecting interrupts at moments unrelated to the tick grid and sampling thread 0 at arbitrary points in its cycle rather than the same one. The injector posts only when nothing is outstanding, so interrupts can never be queued faster than they are serviced. It is confined to the Linux simulation port and no port file changes. The count of windows asked for goes back to ten, on the same reasoning that lowered it to three: ask for what a run can actually reach. Measured over six build configurations of a branch carrying this change, every one reaches ten of ten in under a second, against one to three of three previously with three of the six pinned at the 180 second budget. The budget and the zero window ceiling are untouched, so a run that somehow still falls behind behaves exactly as it does today. The SMP copy deliberately does not get the injector, and its count stays at twenty. It has never been in the slow mode, and injecting interrupts there measurably hurts: 8.9 seconds against 0 to 1 for the same twenty windows, which is the extra interrupt traffic contending across the simulated cores. Only the tx copy has the problem this solves, so only the tx copy changes; the budget and ceiling logic stays identical between them. Verified on this branch: the tx suite passes 105 of 105 with the test at 0.44 seconds, and four direct runs reach ten of ten in 0 to 1 seconds. The SMP suite passes 117 of 117 unmodified, with its own test at 0 to 1 seconds over four runs. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
9e4c57138d |
Normalized the AI disclosure comment to one fixed line per file (#740)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
r52_fvp / r52 (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The per-edit disclosure named the product and model, so every agent and every
model version appended another line. 74 files carried two to four of them, and
the same five products had accumulated 13 spellings -- Copilot against GitHub
Copilot, Claude Sonnet 4.6 against claude-sonnet-4.6, four spellings of Codex.
Twenty assembly lines carried a doubled comment marker, `; //` or `@ //`.
Every file now carries exactly one line, fixed text naming no product:
Portions of this file were generated with AI assistance.
It is written with the comment character that file already uses, so the `;`
and `@` assembly files keep theirs and the doubled markers are gone. Precise
attribution stays on the commit, where the Assisted-by trailer is per-change,
dated and attached to the diff it describes. A header line cannot hold that
record honestly, because the code it names gets rewritten and the line stays.
A file-level flag answers whether; the history answers who.
Comment-only. 455 files, 455 insertions and 574 deletions: every removed line
was a disclosure line, every added line is the fixed text, and no file is left
with zero or with more than one. `scripts/check_ports.sh` passes, including the
reproducibility check that would catch a ports_arch master and its generated
copies drifting apart. Recompiled against dev, every file that builds without a
vendor toolchain gives a byte-identical object: 19 of 19 C files under common,
100 of 100 GNU assembly files, and all 16 assemblable files whose comment
marker changed. The 10 remaining marker changes are ac5 and IAR sources where
`;` already started the comment and only the redundant `//` was removed.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
f5e59a1dd3 |
Completed Windows simulator support and regression coverage (#736)
Completed and stabilized the Win32 and Win64 MSVC simulator ports. - Replaced high-latency host synchronization with bounded scheduler handoffs and critical sections. - Added a high-resolution timer, tick batching, idle fast-forward, shutdown coordination and Win64 extension-pointer support. - Brought the Windows regression tooling and the new thread-transition tests up on CMake, Ninja and the Visual Studio Build Tools. - Extended the SMP teardown diagnostics and corrected 64-bit trace-test handling. This supersedes the historical `win64`, `win32-perf`, `windows-sim-ports` and `windows-sim-ports-completion` branches; no unmerged change from them is missing here. The original Win64 port landed in #529. 1,610 of 1,610 tests pass, across five configurations each: Win32 515, Win64 515, Win64 SMP 580. No external dependency was added, and the existing MSVC warnings in the trace configuration are unchanged. Hardware validation does not apply to host simulator ports. Assisted-by: Codex (gpt-5.6-sol) <codex@openai.com> |
||
|
|
30b3d22adc |
Restored the random stack fill value cleared during thread creation (#732)
Fixes #723 With `TX_ENABLE_STACK_CHECKING` and `TX_ENABLE_RANDOM_NUMBER_STACK_FILLING` both enabled, `_tx_thread_create` picked a random byte, stored it in `tx_thread_stack_fill_value`, filled the stack with it -- and then cleared the whole control block with `TX_MEMSET`. The stack held the pattern while the control block claimed zero, so `TX_THREAD_STACK_CHECK` and `_tx_thread_stack_analyze` compared against the wrong value for the entire life of the thread. The value is now computed into a local and written back after the clear. The clear stays where it is, because the module manager's error checking walks the created list before the control block may be touched. Fixed in `common/src/tx_thread_create.c`, `common_smp/src/tx_thread_create.c` and `txm_module_manager_thread_create.c`, where it additionally left a user-mode module thread's kernel stack filled with zeros. `threadx_thread_stack_fill_value_test` creates sixteen unstarted threads and checks each control block against the pattern in its stack, tolerating a random zero byte without letting that hide the defect. Added to both suites: `ERROR #3` on `dev` in `stack_checking_rand_fill_build`, green with the fix. Five tx configurations at 104 tests, five SMP at 117, and all three sources clean under `-Wall -Wextra` across every combination of the four stack-filling switches. Assisted-by: Copilot (Opus 5) <noreply@github.com> |
||
|
|
3e5a16c95b |
Allowed tx_timer_change to be called from tx_application_define (#730)
Fixes #224 `_txe_timer_change` returned `TX_CALLER_ERROR` for any call at or above `TX_INITIALIZE_IN_PROGRESS`, making `tx_timer_change` the only timer service that could not be called from `tx_application_define` -- while still being allowed from an ISR. The check has no technical basis: `_tx_timer_change` only writes the expiration fields of a timer that is not on an active list, with interrupts disabled. Removed it from both the `common` and `common_smp` copies of `txe_timer_change.c`, with the now-unused includes and the `TX_CALLER_ERROR` line in the header comment. Relaxing an error check is backward compatible. `testcontrol.c` in both suites now calls `tx_timer_change` at initialization, so the timer simple test covers it: `ERROR #30` on `dev`, green with the fix. 116/116 SMP, 103/103 non-SMP. Assisted-by: Copilot (Opus 5) <noreply@github.com> |
||
|
|
1d4a4aeec8 |
Guarded _tx_thread_stack_analyze against inverted stack pointers (#727)
Fixes #460 `TX_ULONG_POINTER_DIF` casts the pointer difference to `ULONG`, so when `tx_thread_stack_highest_ptr` sits below `tx_thread_stack_start` the midpoint wraps to a huge value, the probe lands outside the stack, and the search never converges: the caller hangs or faults. `_tx_thread_stack_analyze` now requires the highest pointer to be strictly above the start of the stack, and bounds the final scan by it. #464 covered the `TX_THREAD_STACK_CHECK` path; this covers direct callers too, as @billlamiework suggested on the issue. Inconsistent pointers still mean a real overflow or a corrupted control block, which remains the application's problem. What changes is that ThreadX reports it through the stack error handler instead of hanging. New cases for an inverted and an equal pointer pair crash the suite without the fix and pass with it. Assisted-by: Copilot (Opus 5) <noreply@github.com> |
||
|
|
cd6a2d9034 |
Put the module manager test's thread on the kernel's created list, so the stand-in kernel matches the one the manager reads (#739)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / riscv (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The host test built a TX_THREAD, gave it an ID and a module instance, and left it off the kernel's created list. A created thread lives on that list, so the thread the test created was not one the kernel would recognise. The created list head and count are now defined for each object type, the thread the test creates goes onto the thread list the way thread create puts it there, and a successful delete takes it off again the way thread delete does. Only the thread list is populated; the others stay empty because nothing here creates an object of those types. No expectation changed. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
83a3e62c28 |
Added a Module Manager host test harness and released a module thread's kernel stack only when it has one (#738)
* Released a module thread's kernel stack only when it has one, so deleting a thread without a manager-allocated stack no longer asks for that memory back The thread delete dispatcher decided whether to release a kernel stack from the calling module's property flags. A user-mode module can be given the address of a thread that carries no kernel stack the manager allocated, and deleting it passed that thread's null stack pointer to _txm_module_manager_object_deallocate(). The pointer is now tested instead of the module's properties. Both thread create paths clear the whole control block before filling it in, so a thread with no manager-allocated kernel stack holds TX_NULL in the field, which makes the test exact rather than an inference from how the module was loaded. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> * Added a Module Manager host test harness and covered the user-mode thread kernel stack lifetime The Module Manager is not in the ThreadX library the test tree links, and its control blocks come from a module port rather than a base port, so nothing in the suite could reach it. The thread_transition directory already describes this technique as the one "the module manager tests in this tree use", but no such tests existed. This adds the directory that comment refers to: a port shim that replaces the three interrupt primitives with host equivalents that count, and a CMake target that compiles the manager sources under test directly against one module port's headers. The dispatch layer is one header of static functions and an unoptimised build emits all of them, so a test that includes it to reach one dispatcher would pull in references to every service the manager can dispatch. The header guards each dispatcher with its own TXM_<SERVICE>_CALL_NOT_USED macro, so the build reads that list out of the header and defines every guard except the ones the test needs. Reading it rather than writing it down means a service added later is excluded without anyone having to remember. The first test asserts the invariant a user-mode module thread has to hold: one logical thread costs the object pool two allocations, the control block the module asks for and the kernel stack the manager takes on its behalf, and deleting the thread must return the pool's available bytes and the module's allocation-list count to exactly what they were. It asserts an equality rather than a bound, because a bound would pass while one of the two was left behind on every cycle, and then runs enough cycles that a leak of one stack per cycle exhausts the pool several times over. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
70741198e9 |
Refused thread delete and reset while an exit transition is in progress (#724)
_tx_thread_shell_entry and _tx_thread_terminate both publish a thread's
terminal state -- TX_COMPLETED or TX_TERMINATED -- and then call that
thread's exit notification callback, before the thread has been detached
from the ready list and before either service has finished with the pointer
it holds to the control block. That terminal state is exactly the state
_tx_thread_delete and _tx_thread_reset accept as authorization to
invalidate or rebuild the control block, and neither service tested whether
the transition producing it had finished.
A callback could therefore delete the thread it was called for -- and then
lawfully recreate it over the same memory, since delete exists to permit
that -- while the scheduler was still linked to the old incarnation.
tx_thread_create zeroes the whole control block and can auto-start the new
one, so the old priority list is left heading at a block whose own priority
field names a different list, with the old priority's map bit set behind
nothing. Alternatively a callback could reset a terminated thread, which
moves it out of the terminal state, and then resume it: the interrupted-
suspension logic in _tx_thread_system_resume refuses to void a suspension
only while the state is still terminal, so with the reset allowed first the
resume clears the suspending flag and restores TX_READY, the outer service
then finds the flag clear and skips the removal, and tx_thread_terminate
returns TX_SUCCESS for a thread that is runnable again.
Interrupt masking does not close the window, because the kernel restores
the prior posture before invoking the callback deliberately. On SMP
TX_RESTORE also releases the global protection, so the target can be
executing on another core while its callback runs -- and a reset there
memsets the stack a live core is running on. On the Linux, Win32 and Win64
host simulation ports the consequence is more immediate than corruption:
TX_THREAD_DELETE_PORT_COMPLETION cancels and joins the host thread backing
the deleted thread, so a callback-side delete of the completing thread
destroys the host thread the callback is running on.
The fix marks the transition and has the two services refuse a marked
target, returning the errors they already document, TX_DELETE_ERROR and
TX_NOT_DONE. The refusal is transient and the same call succeeds once the
transition has completed, so no documented lifecycle is lost; and it is in
the core services rather than the _txe_ wrappers, so disabling error
checking cannot disable it. Refusing the reset is also what closes the
resume path, without touching _tx_thread_system_resume: its existing
terminal-state test is sufficient once nothing can turn the terminal state
into TX_SUSPENDED from inside the window.
tx_thread_suspending is the marker, rather than a new control-block field.
It already means "a suspension is in progress" and is already true across
the callback in the two interruptable paths, so no field is added, the
public structure is unchanged, and sizeof(TX_THREAD) is unchanged --
which matters, because the Module Manager's object handling depends on the
sizes of the control blocks. Widening its lifetime was checked against
every reader rather than assumed. There are four: two in
_tx_thread_system_suspend and two in _tx_thread_system_resume. In every
window this change widens, the state is TX_COMPLETED or TX_TERMINATED, and
both resume readers already refuse to void a suspension for exactly those
two states, so their behaviour is unchanged; and no suspension routine is
called on the target in those windows, so the suspend readers never see
them. No suspension-initiating service can set the marker again inside a
window either: every one of them acts on a thread that is ready or
suspended.
Three sites needed changing beyond the two refusals, and the shape of each
was decided by where the marker can safely be cleared:
- The non-ready branch of _tx_thread_terminate cleared the marker before
the terminated extension and the callback, which is what left them free
to act on a control block the service still had mutex-release
processing to do against. The clear moves to the common tail, after the
last dereference of the target, and becomes the single clear site for
the whole service. In the interruptable ready branch the flag is
already false there, because _tx_thread_system_suspend cleared it when
it detached the thread, so the tail store is a second store of a value
the flag already holds -- cheaper than testing for it, and it keeps one
clear site.
- Under TX_NOT_INTERRUPTABLE neither path set the marker at all, because
that configuration does not use the interruptable suspension path that
sets it. Both now set it before the callback. Interrupts being disabled
there does not help: the callback is reached by a direct call.
- In the TX_NOT_INTERRUPTABLE completion path the marker is cleared
before _tx_thread_system_ni_suspend rather than after it. That call
returns to the scheduler for a thread that is the current thread, which
a completing thread is, and does not come back; clearing afterwards
would leave a normally completed thread marked for ever and therefore
permanently undeletable. Nothing is lost by clearing early there,
because everything from that point to the detachment runs with
interrupts disabled and calls no application code.
The change is the same change twice. All four files are byte-for-byte
identical between common and common_smp at this commit and stay so after
it, so common_smp was written by copying rather than by repeating the
edits. tx_thread_system_suspend.c and tx_thread_system_resume.c, which do
differ between the kernels, are deliberately untouched.
Tests. The in-tree regression test goes to both trees and is byte-for-byte
identical between them. It drives seven scenarios: terminating a ready
non-current target with two peers ready at the same priority, with the
callback attempting the delete and recreating the block if it succeeded;
the same with the callback attempting the reset and then the resume;
terminating a target suspended on a semaphore while owning a mutex, which
is the non-ready branch; natural completion alone at its priority,
including the safe post-completion reset, terminate, delete and recreate at
another priority; self termination; a benign callback, whose notification
count and ordering are unchanged; and the state and boundary cases, where
the new refusal must not fire.
It measures rather than describes. The callback records the published
state, the marker, and the status of every lifecycle service it can reach,
calling the core service as well as the wrapper wherever a refusal is
expected. A snapshot taken under interrupt lockout -- which is the global
SMP protection on an SMP port -- checks that every ready list agrees with
the control blocks it heads, that the priority map agrees with the lists,
and that each execute pointer is a member of the list its own priority
field names. Every walk is bounded, so a corrupted ring costs an assertion
and not a hang, and no test in the suite can hang. Expectations are counted
inside a scenario and gated between scenarios, so a failing kernel reports
how much it failed by without being driven further into its own
corruption.
The consequences are demonstrated from the terminator's context rather than
the completing thread's, which is what makes the pre-fix behaviour an
assertion instead of a wedged simulator. Compiled against the unfixed
sources the test fails 9 of the 24 expectations it reaches in the
uniprocessor tree and 10 of 24 in the SMP tree, and the failures are the
finding: the callback-side delete succeeds, the recreate succeeds, the
consistency snapshot disagrees, the target is neither terminal nor detached
when the service returns, and a peer has left the ready ring the recreated
block hijacked.
The SMP tree gets a second test for the case that needs concurrency. The
victim is excluded to core 1 and spins there without relinquishing while
the controller, excluded to core 0, terminates it, so the callback runs on
one core while the target executes on another. The callback-side reset and
delete must both be refused, and a sentinel written into the unused low end
of the victim's stack must survive -- a reset would have memset the whole
stack before rebuilding the frame. Every wait is bounded, and if the remote
precondition cannot be established the test says so and drops only the
assertions that depend on it rather than reporting a pass it did not earn;
measured over twenty consecutive runs it established the precondition every
time.
TX_NOT_INTERRUPTABLE and TX_DISABLE_ERROR_CHECKING are not among the five
build configurations either tree compiles, and each tree builds the whole
library once per configuration, so neither can be reached from inside the
suites. The lines this change adds under TX_NOT_INTERRUPTABLE are therefore
in no configuration the trees build, and they are where the permanent-
undeletability failure mode lives, so they get their own harness rather
than a compile check: the four sources plus the two error wrappers are
compiled directly into a test executable, once per combination, with
recorders standing behind the scheduler services they call. That is what
makes the marker's value at the moment of detachment directly observable.
It runs 66 expectations under TX_NOT_INTERRUPTABLE, 66 under that with
error checking disabled, 45 under that with notification disabled, and 63
under error checking disabled alone; against the unfixed sources those fail
25, 22, 6 and 20 respectively. The harness lives in the uniprocessor tree
only, because the four sources are identical between the kernels and the
shim replaces the very primitive the SMP port differs in, so a second copy
would compile the same text under the same macros. It is deliberately left
out of the coverage instrumentation, since the same source under different
feature macros has a different line set and merging those would confuse the
union rather than add to it.
Results. Both suites pass in all five configurations with GCC 14: 103 of
103 in the uniprocessor tree, up from 98, and 116 of 116 in the SMP tree,
up from 114. Merged line coverage is 100% in the uniprocessor tree and
5172 of 5183 in the SMP tree, whose eleven uncovered lines are the same
eleven that were uncovered before this change and are in tx_byte_pool_search
and tx_thread_smp_utilities; all four changed files are at 100% line
coverage in both trees, and SMP branch coverage rises from 2819 of 3548 to
2831 of 3556. Cross-compiled with arm-none-eabi-gcc at -Wall -Wextra for
Cortex-M4 against common and for Cortex-A7 SMP against common_smp, all four
files produce no diagnostics at all and an identical warning set to before
the change, under -std=gnu99 and -std=c99 alike -- unlike the module ports,
-std=c99 does not fail on these base ports, and even -Wconversion is clean.
No MISRA deviation is required: explicit comparisons to TX_TRUE, existing
ThreadX types, single-entry and single-exit control flow, no goto, and two
added constant-time tests that change no real-time complexity.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
850a172bac |
Added lazy FPU stacking and QEMU functional tests for RV64 GNU port (#549)
* add lazy FPU stacking to context save/restore Save mstatus/sstatus to stack slot 29 and skip floating-point register save/restore when FS is Off (bits 14:13). This avoids unnecessary FP context work for threads that do not use the FPU. - context_save: check FS in nested and first-level interrupt paths - context_restore: gate FP restore on nested, no-preempt, and preempt paths - use sstatus when TX_RISCV_SMODE is defined, otherwise mstatus * add QEMU virt CMake build and automated test runner Wire the QEMU virt demo into the CMake build system and add a Python/GDB functional test runner, mirroring the risc-v32/gnu port. - Add qemu_virt/CMakeLists.txt to build kernel.elf and register the check-functional-riscv64 target (requires Python3; skipped if absent) - Link kernel.elf with --whole-archive so all ThreadX symbols resolve - Pin _start at 0x80000000 via .text.boot in entry.s and KEEP(*(.text.boot)) in link.lds - Extend demo_threadx.c with fpu_test_val and shorten thread_0 sleep for GDB-driven FPU, timer, and preemption checks - Add test/azrtos_test_tx_gnu_riscv64_qemu.py; verified passing on QEMU virt (FPU, timer interrupt, preemption) * Clean up RV64 PR scope and remove QEMU test integration leftovers Revert accidental RV64 qemu_virt test/CMake integration changes and keep this branch focused on lazy FPU context handling only. Also remove unintended TX_RISCV_SMODE-based mstatus/sstatus save path and align comments/logic to mstatus-only behavior. * Initialize mstatus.FS in RV64 stack build so new threads start with clean FP state Slot 29 was left uninitialized while context restore reads it as an FP-live hint; garbage FS bits could make a new thread inherit the previous thread's floating-point registers. * Add RV64 regression test for the FP state of a newly created thread The test dirties every floating point register, then creates a thread and checks that the stack builder wrote the mstatus slot and that the new thread starts with all floating point registers zeroed. It is registered for RV64 only, since the RV32 stack builder still leaves the slot unwritten. * Completed the RISC-V64 lazy FPU so the restore side matches the save side The lazy FPU save in this branch skips the floating-point stores when mstatus.FS is Off, and records the mstatus it judged that on in frame slot 29. Merged onto current dev, only the save side had that treatment: both restore paths and the scheduler's interrupt-frame path still reloaded the FP registers unconditionally, from slots the save had deliberately left alone. A thread that never touched the FP unit would have had whatever the frame happened to contain loaded into its registers, and FS driven to Dirty on the way out. The guard is added at the three places that consume an interrupt frame: both paths in _tx_thread_context_restore, and _tx_thread_schedule_loop. Each reads slot 29 and skips the FP block when FS was Off, which is the same shape the risc-v32 port already uses. The solicited path is deliberately left alone. _tx_thread_system_return saves the callee-saved FP registers unconditionally, so restoring them unconditionally is consistent; making that pair lazy as well is a separate change, and risc-v32 is the model for it. Verified with QEMU on all five configurations: risc-v64 96 of 96 passing, five configurations, nothing unlinkable risc-v32 95 of 95 passing, five configurations, unchanged functional check-functional-riscv64 passes every check The ninety-sixth test is the one this branch adds. It is load bearing: seeding stack build with FS = Off instead of Initial makes it fail, and restoring the seed makes it pass, so it guards the behaviour the rest of this branch is about rather than passing regardless. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> --------- Co-authored-by: r <r@r> Co-authored-by: Frédéric Desbiens <frederic.desbiens@eclipse-foundation.org> |
||
|
|
6ce8d5cc76 |
Marked every published ThreadX include directory as SYSTEM so applications no longer get warnings from ThreadX headers (#713)
* Marked every published ThreadX include directory as SYSTEM so applications no longer get warnings from ThreadX headers
Commit
|
||
|
|
40db27e843 |
riscv32: spec compliance and regression test fix (#691)
* riscv32: spec compliance and regression test fix Signed-off-by: Akif Ejaz <akifejaz40@gmail.com> * Derived the RISC-V32 frame sizes from the port contract in one place tx_port.h published TX_RISCV_TRAP_FRAME_SIZE for the GNU BSP assembly, but nothing in the port consumed it. Six .S files each rebuilt the same numbers from their own #if, so the interrupt frame size was written out in seven places and the solicited frame size in three. That is the shape that produced the RISC-V64 fault fixed in #708, where the port moved to a padded frame and one copy of the constant did not. The sources now include tx_port.h and take both sizes from it, and no literal frame size remains in the port. TX_RISCV_SOL_FRAME_SIZE joins the contract, since the solicited frame was never published at all. The emitted code is unchanged: 400 and 176 bytes for ILP32D, 128 for soft-float, confirmed by disassembly before and after. Two further corrections: _tx_initialize_low_level carried .global immediately followed by .weak, so the symbol stayed weak and the .global did nothing. Weak is what the port wants, because the example and regression BSPs both provide their own definition, so the stray .global is removed rather than the .weak. Verified with nm that the symbol is still W. The QEMU runner seeded fpu_verified from skip_fpu, so a soft-float run satisfied the FPU gate whether or not the script ever reported the skip. It now starts false and is set only when the skip marker is present, so a run that dies before reaching that point fails instead of passing. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> --------- Signed-off-by: Akif Ejaz <akifejaz40@gmail.com> Co-authored-by: Frédéric Desbiens <frederic.desbiens@eclipse-foundation.org> |
||
|
|
146d57b235 |
Fixed the RISC-V64 trap frame size mismatch in the regression test BSP (#708)
The RISC-V64 port moved its interrupt frame to 528 bytes (65 slots plus 8
bytes of padding, so sp stays 16-byte aligned at a call) and published the
size as TX_RISCV_TRAP_FRAME_SIZE. The port sources were converted to use
it, but the shared regression test BSP was not: its trap_entry still
allocated a hardcoded 65 * REGBYTES, or 520 bytes.
Every interrupt therefore unwound 8 bytes more than it allocated:
trap_entry: addi sp,sp,-520
_tx_thread_context_restore: addi sp,sp,528
On the RISC-V64 regression suite that left 24 of 95 tests failing in the
default configuration, typically as an illegal instruction once execution
reached a corrupted frame. The example BSP under the port directory was
converted with the port and was unaffected, which is why the functional
QEMU test kept passing.
The test BSP now takes both frame sizes from the port it is linked
against, so the two cannot drift apart again. A port that publishes no
contract keeps the historical layout, so the RISC-V32 side is unchanged
until its own port publishes one.
TX_RISCV_TRAP_CALL_FRAME_SIZE is restored to the RISC-V64 tx_port.h. It
was removed as unused when the frame sizes were introduced, but it is
part of the same contract: it is the space a trap entry reserves around a
call into C, and the psABI requires 16 bytes there rather than one
register slot.
Verified on QEMU with every linkable test built, comparing against the
commit before the port change:
before the port change 2 failures out of 95 (both unlinkable)
current dev 24 failures out of 95
with this change 2 failures out of 95 (both unlinkable)
The two remaining failures predate all of this: newlib pulls _impure_ptr
out of R_RISCV_HI20 range for time(), so those two binaries do not link.
RISC-V32 is unchanged at 2 failures across all five configurations.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
9a03838381 |
Stopped the test runners from reporting an incomplete build as test failures (#709)
The cmake test runners drove Ninja with its default keep-going of 1, so the
first failing target ended the build. Every target scheduled after it was
simply absent, and ctest reports a missing binary as a failing test. A
single link error therefore produced a failure count that moved with build
scheduling order rather than with the code.
Measured on the RISC-V64 regression suite, where two targets genuinely
cannot link:
before 79 of 95 test binaries built
after 93 of 95 test binaries built
Fourteen perfectly good binaries were being skipped and counted as
failures. Passing -k 0 lets Ninja finish everything it can; the build still
exits non-zero when a target fails.
Three related problems in the same paths are fixed with it.
A failing configuration used to abort the loop over configurations, so
under set -e the ones after it went unbuilt or untested. The build loops
and the serial test loops now accumulate status and return it at the end,
which is what the parallel test branch already did with wait, and what
cmake_bootstrap.sh already documented for ctest.
Capturing that status removes the set -e protection inside the functions,
so two latent faults become reachable and are closed here. A failed pushd
would have let ctest run in the source tree, where it finds no tests and
reports success; the pushd is now guarded. And ctest's status was
discarded by the popd that follows it, so a configuration with failing
tests returned 0 and was reported as a pass; the status is now carried
past the popd and the summary steps.
Verified on the RISC-V32 suite, which has two genuine failures in each of
its five configurations. Both the serial and the parallel branch now test
all five and exit 8, where the serial branch previously stopped after the
first configuration.
The tx and smp runners are symlinks to scripts/cmake_bootstrap.sh, so they
are covered by the one change there.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
a2800fef16 |
Covered the SMP suspension teardown and the long byte pool search (#677)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
Against the merged SMP coverage report -- every build configuration instrumented and unioned -- sixty-four lines of common_smp/src were uncovered, 5114 of 5178. Fifty-three of them are closed here and the report reads 5167 of 5178. The SMP coverage floor goes from 98 to 99 with it. Thirty-six of the sixty-four were one loop repeated four times: the walk in tx_block_pool_delete, tx_byte_pool_delete, tx_event_flags_delete and tx_queue_delete that releases every thread suspended on the object with TX_DELETED. The suite deletes all four object types after every single test, and that is exactly why the loop never ran. test_control_cleanup in the ThreadX suite deletes the application's objects first and its threads last, so a test that ends with a thread parked on a queue has that thread walked out of it by tx_queue_delete. The SMP suite's cleanup deletes the threads first, deliberately -- it was changed so that no application-owned object is still referenced when the object loops run, which is what stopped a class of teardown hang. The side effect is that all four deletes now run against an empty suspension list. tx_semaphore_delete is the one member of the family that was already covered, because threadx_semaphore_delete_test deletes a busy semaphore on purpose. threadx_object_delete_suspension_test is the same idea for the other four. Two threads suspend on each of a block pool, a byte pool, an event flags group and a queue; the control thread waits on the object's own suspended count through tx_*_info_get rather than on an ordering it cannot guarantee across four cores, deletes the object, and checks both waiters came out with TX_DELETED. Two waiters rather than one so the loop takes its back edge as well as its body, and every wait is bounded in ticks so a suspension that never arrives fails the test instead of hanging it. threadx_trace_entry_update_test and threadx_thread_misaligned_stack_test are ports of the two tests that closed the equivalent gaps in common/src, and close fourteen more lines here: tx_block_allocate 123, 175, 182, 319 and 326, tx_byte_allocate 130, 210, 217, 359 and 366, tx_thread_system_suspend 504 and 560, tx_trace_object_register 221, and tx_thread_create 133. The one substantive change is core confinement. The trace test needs thread 0 to suspend and thread 1 to then release what it waits for; on four cores thread 1 gives the block back before thread 0 has suspended and the update block behind the suspension is never reached, so both threads are excluded from cores 1 to 3. The misaligned stack test needed no such change. threadx_byte_memory_long_search_test closes three of the eleven in tx_byte_pool_search. Lines 264, 267 and 270 are the TX_BYTE_POOL_MULTIPLE_BLOCK_SEARCH limit -- twenty on this port -- where a long search drops and retakes protection so that it cannot lock the other cores out for the whole walk. No byte pool in the suite ever had twenty fragments. This one is filled with small chunks until it refuses another and then has every second chunk released, so the free fragments are never adjacent and cannot be merged, and the request is larger than any of them but smaller than the pool's theoretical total, which is what makes _tx_byte_pool_search walk rather than refuse at the door. The layout is asserted rather than assumed: the test checks the fragment count and checks the probe request really does fail before the workers start, because either would otherwise turn it into a silent no-op. Eleven lines remain and they are not a to-do list. Eight are the delay loop in tx_byte_pool_search that fires when another thread claims the pool inside the window the search opens. The Linux SMP port serialises all four cores on one pthread mutex, so that window is an unlock immediately followed by a lock on that mutex, and glibc hands an uncontended mutex straight back to the thread that just released it: measured over 180,003 windows across three cores, with zero handovers. The shipped test therefore does 250 searches per worker rather than the sixty thousand that probe used, because the twenty-block threshold is crossed by the first search. The other three are in tx_thread_smp_utilities. Line 149 is a range guard placed after the shift it is meant to guard, so reaching it needs a shift by the width of the type; the fix is to move the check above the shift, matching the TX_MAX_PRIORITIES > 32 variant of the same function, and that belongs in its own change. Lines 1073 and 1074 need a mutex owner that is genuinely executing on another core when a waiter suspends, and three shapes were tried without producing one on this port. Measured twice before and twice after, every gcda deleted between runs and 570 of 570 tests passing each time: 5114 of 5178 both times before, 5167 of 5178 both times after. Branch coverage goes from 2768 to 2821 and 2823 of 3548. A floor of 99 needs 5127, so the ratchet lands with forty lines of headroom against a numerator that has been seen moving by two between runs. Assisted-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
f89d65f041 |
Covered the trace entry update paths and the misaligned stack adjustment (#666)
Against the merged coverage report -- every build configuration instrumented and
unioned -- eighteen lines of common/src were uncovered. Seventeen of them were
one missing scenario rather than eighteen separate gaps.
tx_block_allocate, tx_byte_allocate, tx_thread_system_suspend and
tx_thread_system_resume each carry blocks under TX_ENABLE_EVENT_TRACE that go
back and patch a trace entry once the call has done its work, all of the shape:
if (entry_ptr != TX_NULL)
{
if (time_stamp == entry_ptr -> tx_trace_buffer_entry_time_stamp)
entry_ptr comes from _tx_trace_buffer_current_ptr, which stays TX_NULL until
tx_trace_enable is called at run time. Building with TX_ENABLE_EVENT_TRACE is
not enough, and exactly one test in the suite enables tracing --
threadx_trace_basic_test -- which tests the enable API itself and never calls
either allocator. So those blocks sat in the report's denominator and never in
its covered set.
threadx_trace_entry_update_test enables tracing and then drives both allocators
twice each, once on the path that succeeds immediately and once through a
suspension that a second thread satisfies, since each allocator carries one
update block on either side. It then sleeps so that the last runnable thread
suspends with nothing ready to take over: tx_thread_system_suspend lines 345 and
351 are on the branch that sets _tx_thread_execute_ptr to TX_NULL, and the two
allocator suspensions never reach it because the other thread was always ready.
The same test closes tx_trace_object_register's TX_NULL name break by creating a
semaphore with no name. A semaphore and not a thread deliberately: for
TX_TRACE_OBJECT_TYPE_THREAD the register function dereferences the pointer it is
given to read the thread's priority, so that type needs a real TX_THREAD behind
it. threadx_trace_basic_test makes the equivalent call only under
ifndef TX_ENABLE_EVENT_TRACE, against the no-op stub.
threadx_thread_misaligned_stack_test covers the remaining line,
tx_thread_create.c:136, where a stack that does not begin on a ULONG boundary
costs a ULONG of size so that rounding the start up cannot run past the end of
the caller's buffer. Every other test hands tx_thread_create an aligned stack.
That line is compiled only under TX_ENABLE_STACK_CHECKING, so it is absent from
three of the five configurations' reports rather than uncovered in them, and it
was verified under stack_checking_build.
Measured on the merged report, all 480 tests passing and 5 of 5 configurations
green: 4485 of 4503 covered before, 4502 of 4503 after, denominator unchanged.
One line remains, tx_thread_system_resume.c:529, and it is the report's last
flapping line rather than a standing gap -- two clean runs of the same tree gave
4503 of 4503 and 4502 of 4503. Reaching it by construction was tried twice and
failed both times, so it is left alone here. It needs the preempt disable flag
and the system state both clear, and tx_thread_resume raises the preempt disable
flag before calling _tx_thread_system_resume, as do the put and send paths;
creating a higher priority auto-start thread from thread context does not raise
it but does not reach the check either, which a probe showed is executed only
during initialisation, with the system state at TX_INITIALIZE_IN_PROGRESS.
Assisted-by: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
3e85bbd431 |
Instrumented every build configuration and merged their coverage (#665)
Only default_build_coverage carried -fprofile-arcs, because the gate was the build type and it is the only one of five whose name ends in _coverage. The other four build and run all their tests and their coverage was discarded. That is not redundancy thrown away: each configuration selects a different set of TX_ feature macros, so the code the other four compile is absent from the denominator rather than uncovered in it. TX_COVERAGE instruments a build regardless of its name, defaulting to OFF so a single configuration built by hand behaves as before. coverage.sh gains a --merge mode that unions the per-configuration JSON tracefiles, and cmake_bootstrap.sh runs it after the test loop so a local run produces the same merged report CI reads. The template sets TX_COVERAGE for build and test, and coverage_name moves to the merged report. Measured on the ThreadX suite, all 480 tests passing: default_build_coverage 3827 valid 3827 covered disable_notify_callbacks_build 3767 3766 stack_checking_build 3857 3856 stack_checking_rand_fill_build 3862 3861 trace_build 4123 4108 merged 4503 4487 The denominator grows by 676 lines, 17.7%, and the figure moves from 99.97% to 99.64%. The second one is honest, and the drop is the point rather than a regression: the denominator now includes code the old report never counted. The union also contains a file the old report did not contain at all -- tx_thread_stack_error_handler.c compiles only under TX_ENABLE_STACK_CHECKING, so it was not listed at 0%, it was simply absent. 177 files becomes 178. Coverage collection moved out of test() and now runs after the test loop, one configuration at a time. gcov writes its intermediate gcov files into the directory gcovr is rooted at, and coverage.sh roots every configuration at the repository root so filenames come out repo-relative. Five concurrent gcovr processes therefore share one scratch directory and delete each other's output: the first full run of this change passed all 480 tests and produced no report for three of the five configurations. Measured both ways -- two gcovr rooted at the repository root fail concurrently and succeed in sequence. CI would not have caught it, because test_tx.sh sets CTEST_PARALLEL_LEVEL=1 and takes the serial branch. Per-configuration output moved under coverage_report/per_configuration/ and is excluded from the Pages artifact. The deploy job merges the ThreadX and SMP artifacts into one tree and every configuration directory has the same name in both, so left at the top level one suite's would overwrite the other's on the published site. On the SMP suite, an earlier run of this change saw trace_build fail threadx_smp_time_slice_test and then hang, which raised the question of whether -fprofile-arcs perturbs a timing-sensitive test. It does not. Sixteen runs settle it, and the decisive one is that threadx_smp_time_slice_test failed ERROR #31 -- twice in a row under --repeat until-pass:2 -- on an uninstrumented build, in the exact shape CI runs, while three instrumented runs of that shape passed 5 of 5. In the CI shape, CTEST_PARALLEL_LEVEL=1 run.sh test all: TX_COVERAGE=OFF 3 runs 2 green, one ERROR #31 310 s TX_COVERAGE=ON 3 runs 3 green, 5/5 each 325-329 s So the test is a pre-existing flake on dev and instrumenting all five costs about 5% of the suite's wall clock. Separately, and also in both instrumented and uninstrumented builds, run.sh's parallel branch -- what a developer gets typing run.sh test all with no CTEST_PARALLEL_LEVEL -- hangs under its own load, four times in twelve runs. Several SMP tests create 1024 ThreadX threads by construction and the Linux port backs each with a pthread, so five configurations at once put on the order of 5000 threads on the machine. CI sets CTEST_PARALLEL_LEVEL=1 and does not take that branch. Assisted-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
b756220c43 |
Fixed the coverage report's paths and scoping (#664)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The Cobertura XML embedded absolute machine paths, and the flag that looked like it scoped the report to one build configuration was doing nothing at all. Both coverage.sh scripts had the defect; both are fixed here, because the SMP report is published to the same Pages site as the ThreadX one. Paths. -r was the build directory and -f pointed outside it, so gcovr could not express the sources relative to the root and fell back to absolute paths. The result named files as /home/runner/work/threadx/threadx/common/src/... while the <source> element beside them said build/default_build_coverage, so the two halves of the same file disagreed and nothing could map coverage back to the repository. -r is the repository root now and -f an absolute path beneath it, which gives filename="common/src/tx_block_allocate.c". Both must be absolute: -r ../../.. -f common/src produces a report of zero files and exits 0, which is the worst failure mode available here. Scoping. --object-directory does not restrict which gcda files are found -- it tells gcovr how to get from a gcda file back to the compiler's working directory. Pointed at an empty directory it still produced the full 177-file report. That was harmless only by accident, because -r build/$1 constrained the search instead; moving -r to the repository root removes that accident, so the two changes have to land together. Measured, with a second instrumented configuration deliberately made sparser than the first: scoped by the positional search path 3221 of 3827 lines -- the truth no search path, -r at the repo root 3827 of 3827 -- silently merged --object-directory at the sparse tree 3827 of 3827 -- scopes nothing So the search path is load-bearing, and it matters ahead of instrumenting all five configurations: without it each configuration would have reported the union as its own. An empty report is not an error to gcovr -- it warns and exits 0 -- and it carries line-rate="1.0" next to lines-valid="0", so a consumer reads no data at all as fully covered. No coverage threshold can catch that, since an empty report passes any threshold. Hence the explicit assertion that the report has content, which fires with exit 1 on an object directory that exists but is empty, where the old shape returned 177 files and exit 0. Also says out loud that ports/linux/gnu/src is deliberately outside the filter. gcno files exist for it and it is dropped without a word today. Number-neutral, and that was the test. Over the same frozen gcda, changing only the gcovr invocation: ThreadX 3827 of 3827 lines and 1993 of 1994 branches across 177 files, SMP 4739 of 4791 and 2417 of 2430 across 185, before and after alike. Same answers on gcovr 7.0, 8.3 and 8.6, so the change is not wedged to the current pin. End to end through run.sh, 96 of 96 and 110 of 110 pass with the reports written. Assisted-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
eabdb86409 |
Matched gcov to the compiler that produced the coverage data (#658)
gcov reads a data format tied to the compiler that produced it. coverage.sh took whatever gcov was first on PATH, which was fine while the compiler was also whatever was first on PATH. #656 made cmake/linux.cmake honour CC and #657 made a compiler switch actually reconfigure the build, so that assumption no longer holds, and the first person to use the new capability would have hit this. Measured on dev with both of those merged: CC=gcc-14 ./run.sh build default_build_coverage # succeeds ./coverage.sh default_build_coverage # exit 64 gcov says why, if asked directly: tx_block_allocate.c.gcno:version 'B42*', prefer 'B33*' gcovr turns that into "GCOV returncode was 3" and exits 64 through a Python traceback, after the tests have already passed. It reads like a coverage bug rather than a toolchain mismatch, which is the part that would have cost someone an afternoon. gcov is now derived from CC rather than found on PATH, so the caller sets one variable instead of remembering two. GCOV still overrides, for a toolchain that does not follow the gcc/gcov naming, and a derived gcov that does not exist is reported as such instead of surfacing as a traceback. Verified, tx and smp, before and after: CC=gcc-14 was exit 64, now exit 0, 177 files and 1527/3827 lines CC unset exit 0, 177 files and 1527/3827 lines, unchanged CC=gcc-99 exit 1 naming gcov-99 and CC, rather than a traceback GCOV=gcov-14 with CC=gcc-99, exit 0, so the override still wins A mismatched pairing still fails, deliberately: reading a gcc-14 tree with the default gcc-13 gcov is exit 64 as before. Producing a number from mismatched data would be worse than refusing. Assisted-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
e46b1b0787 |
Asked the wait abort test for three windows, and failed a run that reached none (#649)
Four CI runs of the same tree, twenty configuration-runs in total, show this
test's budget being reached far more often than the first green run suggested,
and a pass being reported every time it was:
trace_build 3 of 10 windows in 121 seconds
disable_notify 3 of 10 windows in 121 seconds
default_coverage 4 of 10 windows in 121 seconds
stack_checking 7 of 10 windows in 121 seconds
trace_build 0 of 10 windows in 121 seconds
disable_notify 7 of 10 windows in 121 seconds
stack_checking 3 of 10 windows in 121 seconds
Seven of twenty, and the shortfall message only ever reaches an artifact:
ctest is run with --output-on-failure, so a passing test's output is not in the
job log at all. The suite has been quietly losing most of this test's coverage
in whole configurations and reporting green.
The loop runs in two modes, not one. A window arrives in milliseconds in the
fast mode, and costs between 17 and 40 seconds in the slow one, with nothing in
between across those twenty runs. Ten windows are therefore unreachable inside
any budget worth having: at 40 seconds each that is 400 seconds, and the
unbounded runs measured before any of this took up to 726. Raising the budget
to cover the slow mode would trade a quiet loss of coverage for five
configurations approaching the sixty minute step timeout.
So ask for what a run can reach. Three windows cost 51 to 120 seconds in the
slow mode and under a second in the fast one, and the later hits repeat what
the first ones establish, so what is given up is small. The budget goes to 180
seconds because three windows at the worst rate measured is exactly the 120 it
was, which would have truncated at two.
The count is printed on every run rather than only on a short one. A number
that appears only on shortfall cannot be told apart from a number nobody
recorded.
Reaching the window no times at all is a different matter, and was the worst of
the seven. The check after the loop compares semaphore bookkeeping that a
window has to have touched to mean anything, so a run that reached none of them
compares a counter against the value it was initialised to and reports a pass
having verified nothing. That run now keeps trying to a 300 second ceiling, and
fails if it still has not reached the window. A genuine resonance that holds
for five minutes is worth a failure; the old behaviour was worth nothing.
The SMP copy keeps its count of twenty. It reaches them in under half a second
in all five of its configurations, in all four runs, so the slow mode has never
been observed there and the coverage is free. Both copies get the ceiling and
the unconditional report, so the logic stays identical between them.
Verified locally on all five configurations: the test reaches 3 of 3 in 5 to 14
seconds, and the full suites pass 96 of 96 and 110 of 110 run one test at a
time. With the handler's window made unreachable and the ceiling lowered to 5
seconds, the test stops after 6 seconds, prints the count it reached, and
reports ERROR #8 with the harness recording a failure rather than a pass. With
the count raised past what the budget allows, a run that reaches two windows
still passes, so falling short and reaching nothing stay distinct. The
TX_NOT_INTERRUPTABLE branch, which no configuration in either suite builds, was
compile-checked in both copies with the configurations' own compile commands.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
fe40079353 |
Stopped the thread priority change test leaving core 0 to a finished thread (#647)
The SMP suite has been failing on threadx_thread_priority_change since 30 June.
The test reports SUCCESS and the process then never exits, so ctest kills it at
the thousand second timeout, twice, and the log carries nothing past the result
line. Instrumenting the harness teardown produced the state at the hang:
last stage reached: test_control_return: control thread resume returned
_tx_thread_preempt_disable: 0
core 0: current=thread 0 execute=thread 0
thread test control thread state=0 priority=0 threshold=0 core_control=1
thread thread 0 state=1 priority=0 threshold=0 inherit=0
Core 0 is held by a thread in state 1, TX_COMPLETED, while the control thread
sits in state 0, TX_READY, at priority 0. The Linux SMP scheduler fills a core
only when _tx_thread_current_ptr for it is null, and clears that pointer only
for a thread carrying a deferred preemption, which thread 0 is not. So core 0
can never be handed on, and with TX_THREAD_SMP_ONLY_CORE_0_DEFAULT and
TX_SMP_NOT_POSSIBLE the control thread has nowhere else to go. The scheduler
re-reads the same state every two milliseconds for as long as it is allowed to.
What put thread 0 at priority 0 is the last thing this test does:
thread_0.tx_thread_inherit_priority = 0;
_tx_thread_smp_simple_priority_change(&thread_0, 16);
with the stated intent of reaching the branch where the new priority is below
the inheritance priority. Zero cannot reach that branch, because 16 is not less
than 0. The other branch runs instead, and that branch assigns the inheritance
priority as the thread's priority while the code after it links the thread into
the list for the new priority regardless. Thread 0 therefore came away claiming
priority 0 while living in the priority 16 list.
Both halves of that hurt. Priority 0 ties with the control thread, so resuming
the control thread raised no preemption and left the execute pointer alone. The
mismatch between the recorded priority and the list the thread is linked into
then means that completing thread 0 removes it from a list it was never in,
leaving core 0 pointing at it for good.
Give the inheritance priority a value above the new one, which is what the
branch the comment names actually requires, and put it back to
TX_MAX_PRIORITIES afterwards so nothing downstream reasons about an
inheritance that is not there. Hold protection across the call as well: this is
an internal routine that expects it, and it was being called in the open.
Measured before and after by printing the thread's state at the point the test
reports success. With the inheritance priority at 0 it is priority 0 threshold
0, matching the hang above, on every run. With it above the new priority it is
priority 16 threshold 16, which agrees with the list the thread is linked into,
and resuming the control thread preempts core 0 the ordinary way.
All five SMP configurations pass 110 of 110 at the parallelism CI uses.
The comparison in _tx_thread_smp_simple_priority_change deserves a second look
on its own account. When its else branch runs, the thread's recorded priority
and the list it is linked into disagree by construction. Only this test is
known to reach that branch, by supplying an inheritance priority that cannot
arise in ordinary operation, so nothing here claims a defect in shipped paths.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
83dfbc3534 |
Made a teardown hang in the SMP suite say where it stopped (#646)
The SMP regression suite times out in CI on threadx_thread_priority_change and the log carries nothing that says why. The reason the log is empty is mechanical: test_control_return() opens with fflush(stdout), and that is the last flush before either the test finishes or it wedges. Everything printed after it sits in stdout's block buffer, and ctest discards that buffer when it kills the process at the timeout. So the log ends at the test's own result line no matter where the process actually stopped. That is enough to place the hang, if not to explain it. The failing runs print "SUCCESS!" and then stop, which means every check in the test body ran and passed, and the wedge is somewhere between that flush and exit(). The run of 30 June shows the same signature before any of this test's waits were bounded, so the hang is not the unbounded wait removed earlier, and the message added then for an exhausted cap never appears. Retrying tells us nothing new either: the suite spends 1000 seconds per attempt, twice, to reproduce the same silent timeout. Record how far teardown gets, and bound it. A stage variable is updated at each step from test_control_return() through test_control_cleanup() to exit(), and a watchdog thread armed on entry to test_control_return() reports the last stage reached, the per-core scheduler state, and every thread on the created list, then exits 99. The watchdog covers teardown and not the test body, because the test body has no bounded runtime to hold it to. Several tests here wait on a probabilistic interrupt window: threadx_thread_wait_abort_and_isr_test has been measured between 0.34 and 439 seconds while passing. Teardown is a fixed amount of work that takes milliseconds, so a bound on it cannot turn a slow pass into a failure. The default is 60 seconds, which also means a wedged run now reports in one minute rather than burning the 2000 seconds two 1000-second attempts cost today. The report is written with write() rather than printf() because a wedged thread may be holding the stdio lock, and a watchdog that blocked on that lock would reproduce the silent timeout it exists to replace. For the same reason it reads the ThreadX globals directly and takes no kernel lock; the values may be torn, which is acceptable for a post-mortem and cannot deadlock. One walk in test_control_cleanup() is bounded as well. The loop that steps past the timer thread and the control thread has no terminating condition of its own and spins for good if _tx_thread_created_count and the created list ever disagree, which is one of the shapes the timeout could be taking. It now reports and stops instead. Off by default in the sense that matters: stderr stays empty and stdout keeps its buffering, so output is unchanged on a passing run. TX_TEST_TEARDOWN_TIMEOUT overrides the bound in seconds and zero disables the watchdog; TX_TEST_TEARDOWN_TRACE echoes each stage as it is reached and line-buffers stdout so the surrounding output survives a kill too. Verified against an injected hang at the point the failing runs stop: the watchdog fires, names the stage, and exits 99. The dump is already informative, showing thread 0 left at priority 0 with threshold 0 and inherit 0, the same priority as the control thread, with core 0's execute pointer still on it. All five SMP configurations pass 110 of 110 at the parallelism CI uses, in both quiet and trace modes, and the suite runtime is unchanged. Only the SMP harness is instrumented. The non-SMP suite has not shown this hang. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a5483f0773 |
Bounded the wait for the delayed suspension window (#645)
threadx_thread_delayed_suspension_test waits for an interrupt to land while
thread 2 is part way through suspending, and waits for it with no bound:
while(delayed_suspend_set == 0)
{
tx_thread_wait_abort(&thread_2);
tx_thread_relinquish();
}
How long that takes depends on the build to a degree that is easy to miss. The
loop finishes in between a tenth of a second and three seconds in four of the
five ThreadX configurations. In trace_build it took 490 seconds, which was 40
percent of the whole ThreadX suite and more than every other test in that
configuration put together.
This is the third test in these suites built the same way, after
threadx_thread_priority_change and threadx_thread_wait_abort_and_isr_test: spin
until an interrupt happens to land in a narrow window, with nothing to stop the
spin if it does not. The other two have been given bounds already.
Give this one a wall clock budget too, for the same reason as the last: a tick
is delivered only when the port's timer thread runs, so the tick clock falls
behind real time under load or instrumentation, and instrumentation is exactly
what trace_build turns on.
The check after the loop needs care that the other two did not. It compares
thread_2_counter against thread_2_counter_capture, and the capture is taken
inside the interrupt handler at the moment the window is hit. Leaving that check
in place after a run that never reached the window would compare a live counter
against the zero it was initialised to and report a defect that is not there. So
the check is skipped when the window was not reached, and the run says so.
Reaching the window still exercises it exactly as before.
Verified both ways in trace_build, which is the configuration that was slow: the
window is reached in 8 seconds here and the test passes as it always did, and
with the budget forced to zero the test reports that the window was not reached
and passes without the dependent check firing.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
4d9ce41845 |
Gave the wait abort ISR test a budget instead of an open-ended wait (#644)
threadx_thread_wait_abort_and_isr_test waits for an interrupt to land while the
preempt disable flag is set, and waits for it ten times, twenty in the SMP copy,
with no bound on how long that takes. The handler in the same file already says
what can go wrong:
It is possible for this test to get into a resonance condition in which
the ISR never occurs while preemption is disabled
and perturbs its own duration to break out of it. That helps but guarantees
nothing, and if the resonance holds, the loop does not end.
It is also, by a wide margin, the most expensive thing in either suite. Run one
test at a time in CI it took between 148 and 726 seconds per configuration:
1936 seconds of the ThreadX suite's 2209, against 273 seconds for the other
ninety five tests together. Nothing else in the suite is within two orders of
magnitude of it.
The budget is in wall clock seconds, not ticks. That distinction turned out to
matter. A tick is delivered only when the port's timer thread gets to run, so
the simulated clock falls behind real time under load or under coverage
instrumentation, and never makes the loss up. A first attempt bounded the wait
at 20000 ticks, nominally 200 seconds, and failed to stop a run that took 726
seconds, because fewer than 20000 ticks had gone by. tx_time_get() cannot bound
elapsed time here; time() can.
A run that falls short says how many windows it reached rather than going quiet,
and the check after the loop is untouched. That check compares semaphore
bookkeeping which holds whatever number of windows were hit, so it still means
exactly what it did before. Hitting the race a few times rather than ten is a
smaller loss than it looks: the value is in reaching the window at all, and the
later hits repeat what the first ones established.
Verified by forcing the budget to 3 seconds, where the test stops after 3.14
seconds of wall clock and reports reaching 0 of 10 windows, with the following
check intact. At 120 seconds both suites pass every configuration run serially.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
b3486f9b46 |
Bounded the wait in the thread priority change regression tests (#640)
Both copies of this test, the SMP one and the non-SMP one, install an interrupt
handler and then spin until it clears a flag:
test_isr_dispatch = test_isr;
do
{
...
} while (test_isr_dispatch);
The handler clears that flag only on a narrow window: thread 3 at priority 6,
ready, and not yet at the head of its priority list, which exists only part way
through a priority change. If an interrupt never lands inside that window, the
loop never ends.
That is what has been failing in CI. The SMP suite has been red since 30 June,
and this test times out in three of the last four failing runs, always in a
stack-checking configuration. The evidence that it is a hang rather than slow
work: the test carries no per-test timeout property, so ctest's --timeout 1000
applies, and locally the test finishes in 0.12 seconds with a worst case of 0.29
over thirty runs. Nothing turns that into more than a thousand seconds. It also
survived --repeat until-pass:2, so it hung twice in succession.
The TX_NOT_INTERRUPTABLE path in the same handler already stops after a fixed
amount of work. Only the interruptable path, which is the one the failing
configuration uses, had no protection.
Cap the loop and clear the handler on the way out. When the window is not reached
the test says so and still passes: not reaching it is a gap in what this run
covered, not a fault in the code under test, and failing would report a defect
that does not exist. The counters are left alone so the checks that follow keep
their previous meaning.
The cap is 100000 attempts. An exhausted cap takes about 60 seconds, measured,
and a successful run takes 0.12 seconds, which puts the usual cost around two
hundred attempts and leaves the cap roughly two orders of magnitude clear of it.
That is wide enough not to lose coverage on a slower machine, while replacing a
timeout that says nothing with a message that says what happened.
Not reproducible here, which fits the diagnosis rather than contradicting it: on
sixteen cores the window is hit almost at once. Thirty sequential runs, two
hundred at parallelism thirty-two, sixty pinned to two CPUs, forty pinned to one,
and three full-suite passes at the parallelism CI uses all came back clean. The
defect is the reliance on the window, not any particular machine.
Verified with the window deliberately made unreachable: before this change the
test runs until it is killed, and after it exits in about a minute reporting that
the window was not reached. The full SMP suite passes 110 of 110 at CI's
parallelism.
|
||
|
|
2f6945475b |
Stopped pthread_self() faulting when the caller is not a pthread (#627)
Nothing prevents an application mixing tx_thread_create() with the POSIX layer,
and a thread created that way has no POSIX control block. posix_thread2tcb()
returns NULL for it, which posix_thread2tid() then read through:
p_tcb = posix_thread2tcb(thread_ptr);
thread_ID = p_tcb->pthreadID;
pthread_self() went on to compound it, reading the signal fields of a POSIX_TCB
out of a thread that is only a TX_THREAD:
if (((POSIX_TCB *) thread_ptr) -> signals.signal_handler)
The first is a null dereference and the second runs off the end of the control
block into whatever the linker put there. Under qemu-system-riscv32 the first one
lands first: mcause=0x5, a load access fault, with mtval=0xb4 for the offset of
pthreadID.
Have posix_thread2tid() report zero for a thread with no control block, which is
what px_pth_join.c already does for the same call, and have pthread_self() skip
the signal check unless the ID says the caller really is a pthread. Zero cannot
collide with a real ID because px_pth_create.c uses the address of the control
block as the ID.
This also covers the case where there is no current thread at all, from an ISR or
before the scheduler starts: tx_thread_identify() returns NULL, and the same
zero comes back instead of a fault.
Add posix_pthread_self_test, which asks both kinds of thread for their ID: a
pthread, which has to report what pthread_create() returned, and a plain ThreadX
thread, which has to report zero. Reverting either half of the fix turns the test
into the load access fault above.
Verified with riscv64-unknown-elf and qemu-system-riscv32: 4 tests across the
default build, 4 of 4 passing.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
f184531e7f |
Added the POSIX compatibility layer to the CMake build, with regression tests (#626)
* Added the POSIX compatibility layer to the CMake build
Nothing in the repository built the POSIX layer. The FreeRTOS layer next to it
has had a target since the CMake build was introduced, so the POSIX one was the
odd one out, and 106 source files went unbuilt by any target, on any
architecture.
Add a posix-threadx target, following the FreeRTOS layer's shape: a static
library, EXCLUDE_FROM_ALL so the default build is unchanged, linking threadx and
publishing its own directory as a PUBLIC include path. The sources live in their
own CMakeLists.txt rather than the top-level file, as common/ does, because there
are 106 of them. The seven posix_*.c files in the same directory are a demo and
standalone signal tests, each with its own entry point, so they stay out of the
library.
The layer does not suit every configuration, and the target is only offered where
it can work:
- Hosted simulation ports (linux, win32, win64) build against a C library that
already provides errno.h, pthread.h and the rest. The layer replaces those.
Its pthread.h even uses _PTHREAD_H, the same include guard as glibc's, so its
declarations are skipped wholesale and the build fails on missing types.
Neutralising that guard only exposes the real problem: 69 conflicting
definitions in a single translation unit, for time_t, struct timespec,
sigset_t, pthread_t, pthread_mutex_t, sem_t and more. Both the layer and the
C library implement POSIX, and only one of them can define those names. The
linux port also emulates threads by calling the C library's pthread_create
and sem_wait, which the layer exports itself, so linking the two would divert
the port into the layer that sits on top of it.
- SMP builds. px_int.h declares _tx_thread_current_ptr as a plain pointer,
which is a per-core array under SMP, and the layer tracks no current core.
Building the layer for the first time exposed one portability defect worth
fixing rather than working around. tx_posix.h defined ssize_t as INT, with a
comment conceding it should come from <sys/types.h>. That is correct only where
the C library agrees: on AArch64 newlib makes ssize_t 64 bits, and every
translation unit that reached a library header failed to compile. Defer to the
library when it has declared the type, keyed on the _*_DECLARED guards newlib
uses, and do the same for mode_t, which had the same problem waiting. Where no
library declaration exists the previous definitions still apply, so the 32-bit
targets that did build are unaffected.
Verified by building posix-threadx for arm9, arm11, cortex_m0, cortex_m3,
cortex_m4, cortex_m7, cortex_m33, cortex_m55, cortex_m85, cortex_a7, cortex_a9,
cortex_r4, cortex_r5, cortex_a34, cortex_a53 and cortex_a55 with
arm-gnu-toolchain-14.3.rel1, and for risc-v32 and risc-v64 with
riscv64-unknown-elf: 18 of 18, 106 objects each. cortex_a78 has no non-SMP port
and fails to configure with or without this change. Linking the result against
libthreadx.a leaves only tx_application_define, _tx_initialize_low_level, the
optional execution profile hooks, and memset and strlen unresolved, all of which
the application or its C library supplies. The default build still produces
libthreadx.a and no POSIX library.
Compiling is not the same as working, and on the 64-bit targets in that list it
is not enough. The layer carries a message by putting the address of a private
buffer into the queue, and ULONG is 32 bits on every port, so that address only
fits when TX_64_BIT is defined. Without it px_mq_send.c truncates the pointer
and px_mq_receive.c casts the truncated value back, which GCC reports as nothing
worse than a -Wpointer-to-int-cast warning. Defining TX_64_BIT is not a remedy
either: tx_api.h then reaches for the extension pointer macros, which need
tx_thread_extension_ptr in the thread control block, and outside ports_smp and
ports/linux no port declares it. So the target builds everywhere, but the
message queues are only sound on the 32-bit ports. That is pre-existing, it is
not made worse here, and it is left for a change of its own.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Added regression tests for the POSIX compatibility layer
The POSIX layer had no tests. The seven posix_*.c programs shipped beside it are
demos: they print nothing, report no result and end in infinite loops, so they
tell a person watching a debugger something and an automated run nothing.
Add a suite under test/posix, laid out like the FreeRTOS one and driven the same
way, with scripts/build_posix.sh and scripts/test_posix.sh over a run.sh that
takes the same arguments as its RISC-V counterpart.
The tests run on emulated hardware because they have nowhere else to go. The
layer replaces the C library's POSIX headers and exports the same symbols the
linux port calls to emulate threads, so a host build is not available to it. The
RISC-V QEMU harness that the ThreadX suite already uses is, and this suite reuses
its BSP and testcontrol.c rather than growing copies of them.
Three tests to start:
- posix_mq_basic_test sends a message through a queue and checks the contents
and priority survive the round trip.
- posix_mq_send_abort_test covers the leak fixed in #624, by filling a queue,
blocking a sender on it, aborting the wait and watching the queue's byte
pool. Reverting the fix makes it fail on the pool check, so it measures what
it claims to.
- posix_pthread_basic_test covers pthread creation, a mutex, a semaphore
handoff, pthread_self and collecting an exit value through pthread_join.
The queue's pool is sized (mq_maxmsg + 1) * (mq_msgsize + 11), which leaves room
for about one message beyond a full queue, so the abort test uses small messages
and a shallow queue. With a larger message the first leaked buffer exhausts the
pool, tx_byte_allocate fails, and the sender disappears into the endless loop in
posix_internal_error() instead of reporting anything. Sizing it this way keeps
the failure legible as a pool measurement rather than a timeout.
riscv32 only, and the reason is the layer rather than the harness. The layer puts
the address of a message buffer into the queue, ULONG is 32 bits on every port,
and a 64-bit address only fits there when TX_64_BIT is defined. Defining it makes
tx_api.h use the extension pointer macros, which need tx_thread_extension_ptr in
the thread control block, and no port outside ports_smp and ports/linux declares
it. Configuring for risc-v64 stops with that explanation rather than building
something that would corrupt a pointer at runtime.
Verified with riscv64-unknown-elf and qemu-system-riscv32: 3 tests across
default_build, disable_notify_callbacks_build, stack_checking_build and
trace_build, 12 of 12 passing.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
cfed1d8095 |
Fixed the kernel object leaks on the static creation error paths (#584)
xQueueCreateStatic() and xTaskCreateStatic() take their storage from the caller, so neither leaks memory, but both create ThreadX objects and both return NULL when a later step fails. The caller is left without a handle and cannot call the matching delete function, so any object already created stays registered in the kernel, pointing into a caller buffer that the application is now free to reuse or discard. Three paths were affected. xQueueCreateStatic() abandoned the read semaphore when the write semaphore could not be created. xTaskCreateStatic() abandoned the notification semaphore when the thread could not be created, and abandoned both the semaphore and the thread when the thread could not be resumed. Delete what was already created before returning on each of them. The resume path terminates the thread before deleting it, since a thread created with TX_DONT_START is suspended rather than terminated, which is the same order the idle task uses when it reaps a deleted task. Extend the regression suite to cover all three paths, and add thread resume to the set of entry points the harness can force to fail. Each static failure case now uses its own control block, so a future regression on one path cannot carry damage into the next case and report misleading counts there. Verified against the layer as it stands on dev, where the three new checks fail with the objects left behind, and against the fixed layer, where the suite passes. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
f3df5f9dde |
Added a regression suite for the FreeRTOS compatibility layer (#583)
* Fixed the resource leaks on the xQueueCreate error paths xQueueCreate() allocated the queue descriptor and its backing memory, then created two ThreadX semaphores, and returned NULL on either semaphore failure without releasing anything. Since no handle reached the caller, vQueueDelete() could not be used to recover, so both allocations were lost. A failure on the second semaphore additionally abandoned the read semaphore it had already created, leaving a live ThreadX control block inside freed memory. Release the backing memory and the descriptor on both paths, and delete the read semaphore before returning when the write semaphore cannot be created. This is the teardown order vQueueDelete() already uses, and it matches the cleanup xTaskCreate() performs on its own error paths. Verified with a fault injection harness that intercepts the ThreadX byte pool and semaphore entry points to force tx_semaphore_create() to fail on a chosen call. On a read semaphore failure the layer previously performed 2 allocations and 0 releases, and on a write semaphore failure 2 allocations, 0 releases and 0 semaphore deletions. It now performs 2 releases in both cases and deletes the read semaphore in the second, with the byte pool restored to its prior state. Fixes https://github.com/eclipse-threadx/threadx/issues/570 Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> * Added a regression suite for the FreeRTOS compatibility layer The compatibility layer had no tests in this repository, which is awkward for its creation functions in particular. Each of them takes one or two byte pool allocations for its bookkeeping and then creates ThreadX kernel objects, and each returns NULL when a kernel object cannot be created. The caller is left without a handle, so it cannot call the matching delete function, and anything the layer failed to release is gone until the system restarts. A leaking version and a correct version are indistinguishable from the outside, which is how the leak in issue 570 went unnoticed. Add a suite that counts what the layer takes and gives back. A test asks the harness to fail a chosen kernel creation call, then checks the number of byte pool allocations, releases, object creations and object deletions performed. The ThreadX entry points are intercepted with the linker's --wrap so that tx_freertos.c is compiled exactly as it ships, with no test hooks in it. Note that tx_api.h maps the public API onto the error checking entry points, so the _txe_ symbols are the ones wrapped. Coverage is the creation and teardown paths of queues, tasks, semaphores, mutexes, event groups and timers, including a regression test for the two paths fixed for issue 570. The suite follows the layout of the existing ThreadX and SMP suites, is registered with ctest, and runs in CI through the shared regression template. It is built 32 bit because the Linux port defines ULONG as unsigned int on x86_64 while the layer passes pointers through ULONG arguments, so a 64 bit build truncates them. It is Linux only because --wrap has no MSVC equivalent, and the CMake configuration says so rather than failing at link time. Validated by building the suite against the layer as it stands before the issue 570 fix, where the two expected checks fail with the leaked counts, and against the fixed layer, where all three tests pass. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
b880ffeada |
Fixed SMP execution profile total getters (#553)
Updated the SMP execution profile aggregate getters to copy each core's total into the matching output array element. Added a focused regression test for thread, ISR, and idle total getters. Co-authored-by: Codex <codex@openai.com> |
||
|
|
730b61874b |
Added copyright headers to files missing them
Applied the standard MIT license header to all project-owned C, header, assembly, shell, and Python files that were missing a copyright notice. Third-party, toolchain startup, and auto-generated files were excluded. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
7486de06c8 |
Refactored, consolidated, and cleaned up RV32/RV64 ports (#536)
risc-v: refactor, consolidate, and fix RV32/RV64 ports Consolidates the RISC-V 32-bit and 64-bit GNU/Clang port sources, fixes two pre-existing assembly bugs discovered during testing, and hardens the build infrastructure for both the regression suite and the CORE-V MCU example. --- Port consolidation (RV32 GNU + Clang) --- - Delete ports/risc-v32/clang/src/ (8 .S files had no Clang-specific directives; diverged from GNU only due to missing bug fixes). The Clang port CMakeLists.txt now compiles from ../gnu/src/. - Change .global -> .weak for _tx_initialize_low_level in gnu/src/ to allow BSP-level override without a linker conflict (adopted from Clang port). - Create ports/risc-v32/common/tx_port_riscv32_common.h with all definitions shared between GNU and Clang ports. Reduce both tx_port.h files to thin wrappers. - Add a prominent comment in risc-v64/gnu/inc/tx_port.h explaining why LONG/ULONG are intentionally 32-bit on RV64 (ThreadX ABI requirement, mirrors win64/MSVC LLP64). --- Shared CMake helper --- - Add cmake/threadx_riscv_port.cmake with threadx_add_riscv_port(). All three port CMakeLists.txt files are reduced to ~8 lines each. Include path is relative to CMAKE_CURRENT_LIST_DIR so the helper works whether ports are built standalone or as a subdirectory of the test framework. --- Shared example-build drivers --- - Create canonical driver files under ports/risc-v_common/: inc/csr.h (uintptr_t-based; portable RV32 + RV64) example_build/plic/ (plic.c, plic.h) example_build/uart/ (uart_qemu_ns16550.c/h; static inline putc_nolock) example_build/trap/ (trap_qemu.c; XLEN-portable mcause constants) - Replace per-example copies with symlinks in all qemu_virt and cva6_ariane example directories. - Fix OS_IS_INTERRUPT typo (was OS_IS_INTERUPT) in shared trap_qemu.c. - Gate print_hex() behind TX_RISCV_TRAP_DEBUG. --- Bug fixes in RV32 assembly --- tx_thread_schedule.S: - Solicited-return FP path: reload t0 from the mepc stack slot before csrw mepc, t0. After the FP restore block, t0 held the fcsr value (0 for new threads), which caused mepc = 0 and an immediate instruction-address fault on the first context switch. - Same path: reload t0 from the mstatus stack slot before csrw mstatus, t0 to avoid writing the stale fcsr value into mstatus. tx_thread_system_return.S: - FP callee-saved registers were saved unconditionally before the mstatus.FS check, causing an illegal instruction trap (mcause=0x2) when a thread with FS=Off (lazy FPU, thread has never used FP) voluntarily yielded. - Apply the same FS guard pattern used in tx_thread_context_save.S: read mstatus first, isolate FS[1:0], and skip fsw/fsd if FS == Off. Both bugs were pre-existing on origin/dev and are unrelated to the consolidation changes. --- RV64 64-bit pointer compatibility --- - Add TX_TIMER_INTERNAL_EXTENSION, TX_THREAD_CREATE_TIMEOUT_SETUP, and TX_THREAD_TIMEOUT_POINTER_SETUP to risc-v64/gnu/inc/tx_port.h to store the thread timeout pointer in a VOID * extension field rather than truncating it into a 32-bit ULONG. Mirrors the win64 port pattern. - Define TX_TIMER_EXTENSION_PTR_DEFINED as a portable sentinel. - Update threadx_thread_basic_execution_test.c guard from #if defined(_WIN64) to #if defined(_WIN64) || defined(TX_TIMER_EXTENSION_PTR_DEFINED). - Disable -Wconversion for the RV64 test build: ULONG = unsigned int (32-bit) is intentional for ThreadX ABI but triggers spurious warnings when sizeof() (8 bytes on RV64) appears in arithmetic with ULONG in common/src/. --- Regression suite cmake fixes --- test/tx/cmake/riscv/regression/CMakeLists.txt: - Build testcontrol_weak_defaults.c as a separate OBJECT library and include it in every test executable via $<TARGET_OBJECTS:>. GNU ld does not extract objects from a static archive to satisfy weak symbols, so bundling it in test_utility was insufficient for the standalone threadx_initialize_kernel_setup_test. test/tx/cmake/regression/CMakeLists.txt, test/smp/cmake/regression/CMakeLists.txt: - Same fix applied to the Linux and SMP regression builds. The symbols abort_all_threads_suspended_on_mutex, suspend_lowest_priority, and abort_and_resume_byte_allocating_thread were introduced by the win64 merge and left the standalone test unlinkable. --- CORE-V MCU toolchain and build fixes --- cmake/riscv64-gcc-rv32imc.cmake: - Resolve riscv64-unknown-elf-gcc via PATH so the riscv-collab toolchain in /opt/riscv/bin is preferred when it appears first. ports/risc-v32/gnu/example_build/core_v_mcu/bsp/clz.c (new): - The riscv-collab toolchain is built without rv32 multilib, so its libgcc does not define __clzsi2 (the helper emitted for __builtin_clz() in fll.c). Add a weak __clzsi2 fallback so the build is self-contained with any riscv64-unknown-elf toolchain. The weak attribute yields to a libgcc-provided strong symbol when the Ubuntu multilib package is used. core_v_mcu/CMakeLists.txt: - Add bsp/clz.c to sources. - Reference CMAKE_TOOLCHAIN_FILE via message(STATUS) to suppress the false- positive "Manually-specified variables were not used by the project" CMake warning and to show the active toolchain at configure time. --- Housekeeping --- - Rename azrtos_test_* -> threadx_test_* (eliminate Azure RTOS branding). - Add RV64 QEMU CI test script: ports/risc-v64/gnu/example_build/qemu_virt/test/ threadx_test_tx_gnu_riscv64_qemu.py - Normalize entry.s -> entry.S in all 4 example directories. - .gitignore: exclude build_m7/ and .codex local artifacts. - CI: comment out the riscv regression workflow job and remove it from the deploy job's needs list (preserved in-place for easy re-enablement). --- Verified --- - 95/95 RV32 regression tests pass (QEMU virt) - 95/95 RV64 regression tests pass (QEMU virt) - All 5 Linux build configurations build cleanly (default_build_coverage, disable_notify_callbacks_build, stack_checking_build, stack_checking_rand_fill_build, trace_build) - CORE-V MCU example_build links cleanly with /opt/riscv toolchain Co-authored-by: Copilot 223556219+Copilot@users.noreply.github.com |
||
|
|
2c16114a45 |
Added win64 ports of ThreadX and ThreadX SMP (#529)
Windows x64 port and regression suite This PR adds the Windows x64 (Win64) simulation port for both the standalone and SMP variants of ThreadX, along with the full CMake build and test infrastructure needed to run the regression suite on Windows. New ports Win64 standalone (ports/win64/vs_2022): self-contained Windows simulation port using Win32 threading primitives as virtual cores. Includes CMake integration, build/test scripts, and MSVC project files. Win64 SMP (ports/win64_smp/vs_2022): multi-core Windows simulation port. Supports up to 4 virtual cores backed by Windows host threads. Scheduler and timer improvements The initial port used coarse polling and synchronous SuspendThread/ResumeThread pairs throughout the scheduler hot path. Several rounds of optimization reduced the SMP regression suite runtime from ~150 s to ~78 s (-48%), with no regressions: - Replaced scheduler polling with an event-driven wake path; switched the simulated timer to one-shot rearming to eliminate catch-up ticks. - Skip SuspendThread when _tx_thread_preempt_disable != 0 (new suspension type 3) -- the primary optimization, yielding up to 7.9x speedup on preemption-heavy tests. - Skip SuspendThread when a thread is spinning on the Win32 critical section (suspension type 4), and fix a stale-TLS bug in _tx_win32_critical_section_obtain that could stamp mutex_access on the wrong virtual core. - Added a 2 ms scheduler event timeout (matching the Linux SMP port) to prevent stalls on any missed SetEvent. - Enabled high-resolution waitable timers (SetWaitableTimerEx) for accurate 100 Hz tick cadence. - Increased TX_WIN32_CONTENTION_PAUSE_COUNT from 64 to 256 to reduce SwitchToThread overhead under heavy CS contention. Build and test infrastructure - Hardened the Windows build wrapper (scripts/build_tx.ps1): invoke Ninja directly for Ninja build trees, fix timeout detection, add a default build timeout, and limit fallback replay to real timeout cases. - Added -Clean support to Windows test scripts to remove stale CTest state before each run. - Skip Visual Studio DevShell re-entry when the active MSVC environment already matches the requested architecture. - Fixed scripts/build_tx.sh (Linux) regression source generation: replaced brittle exact-string insertion with line-based matching so the interrupt dispatcher hook is inserted reliably for both simulator ports. Test suite updates - Introduced test/tx/regression/threadx_test_port.h with portable macros (TX_TEST_POINTER_WORD, TX_TEST_STORE_POINTER) for storing pointers in test arrays on 64-bit targets where ULONG remains 32-bit. - Adjusted pool-capacity and pointer-storage patterns in regression tests to use ALIGN_TYPE-sized slots, making the suite correct on 64-bit hosts. - Restored stricter event flag, sleep, and timer expectations now that port-level fixes make prior Windows accommodations unnecessary. - Tightened SMP watchdog and clean-build timeout defaults. Version metadata Updated Win32, Win64, and Win64 SMP port version strings to 6.5.1.202602. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: Codex (gpt 5.5) <codex@openai.com> |
||
|
|
c1e3678797 |
Add QEMU based CI regression test infra for RV32 and RV64 (#526)
Added a QEMU virt-machine BSP and CTest infrastructure to run the ThreadX regression suite on both RISC-V 32-bit and 64-bit targets in CI. New components: - BSP (entry, trap, PLIC, CLINT timer, UART, linker script) targeting QEMU virt machine for RV32 and RV64 - CMake build system with Ninja, supporting multiple build configs - CI scripts: install_riscv.sh (toolchain + QEMU), build_tx_riscv.sh, test_tx_riscv.sh - GitHub Actions workflow job for RISC-V regression gating Port fixes: - RV32 tx_thread_context_restore.S: set MPIE alongside MPP (0x1800 → 0x1880) so mret re-enables interrupts - RV32/RV64 tx_port.h: add TX_REGRESSION_TEST extension macros needed by the test harness - RV32/RV64 example_build scripts: add compile and QEMU launch steps Regression test portability fixes: - Block memory tests: increase pool sizes (320 → 340) to accommodate larger RISC-V block-header alignment - Byte memory test: replace hardcoded offsets with BYTE_POOL_OVERHEAD macro for portable pool-size computation - Event flag timeout test: make counter tolerance unconditional, removing linux-only guard Signed-off-by: Akif Ejaz <akif.ejaz@10xengineers.ai> |
||
|
|
3c4d20285f | Corrected typos in two test filenames (#527) | ||
|
|
f81d1c9eb5 | Merge branch 'dev' into typo | ||
|
|
c3259a2160 |
Updated copyright headers and version number constants (#509)
* Updated version number constants * Removed revision history from all files * Added Eclipse ThreadX contributors' copyright header |
||
|
|
1f59529034 | Added missing QUEUE_MESSAGE_MAX_SIZE test for SMP. | ||
|
|
87e5110346 |
Make queue max message size configurable
Summary ------- This commit fixes the issue #424 Details -------- - Add a the configuration option TX_QUEUE_MESSAGE_MAX_SIZE in the tx_api.h with default value set to TX_ULONG_16 to keep backword compatibility. - Update the txe_queue_create() to check on TX_QUEUE_MESSAGE_MAX_SIZE rather than TX_ULONG_16 as max message size. - Add a new unitary test to cover the new change. |
||
|
|
898d0ebfde | Updated CMake minimal version to 3.13 |