mirror of
https://github.com/eclipse-threadx/threadx.git
synced 2026-10-06 06:59:08 +08:00
508af549dae1c269088b8e02c45670bf87de33b1
160
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
508af549da |
Fixed the incorrect loop bound constant in the IAR file lock support (#695)
The IAR multithreaded library support code allocates its file lock mutexes from an array of _MAX_FLOCK entries, but the wrap-around check and the exhaustion check in __iar_file_Mtxinit() both compared against _MAX_LOCK, the bound of the unrelated system lock mutex array. When _MAX_FLOCK is greater than _MAX_LOCK, the free mutex index wrapped early and the exhaustion check reported failure while free entries remained, so *m was set to TX_NULL and the application faulted the first time a file lock was taken. When _MAX_FLOCK is smaller than _MAX_LOCK, the free mutex index was allowed to run past the end of __tx_iar_file_lock_mutexes and the exhaustion check could never fire. Corrected all four comparisons in each of the 27 copies of tx_iar.c. Fixes #444 Assisted-by: Copilot (Opus 5) <noreply@github.com> |
||
|
|
c469a7756b |
Fixed the missing immediate prefix on MOV in Cortex-M schedulers (#693)
The BASEPRI-masking path in tx_thread_schedule wrote "MOV r0, 0" rather than "MOV r0, #0". UAL requires the "#" prefix on an immediate operand. GNU as and the LLVM-based assemblers accept the unprefixed form and emit the intended encoding, but stricter assemblers reject it outright, so the affected ports could not be built with those toolchains. The GNU and AC6 sources had already been corrected; this brings the IAR and AC5 sources into line. Verified that both spellings assemble to the same Thumb-2 encoding (f04f 0000), so this is a source-correctness fix with no change in generated code or runtime behaviour. Covers 31 occurrences across the Cortex-M3, M4, M7, M33, M52, M55 and M85 ports, their module manager counterparts, and the shared ARMv7-M and ARMv8-M architecture sources. Fixes #461 Assisted-by: Copilot (Opus 5) <noreply@github.com> |
||
|
|
13c8c768c7 |
Added the boot-at-EL1 option to the S32Z280 entry path (#690)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
gcc_check / gnu (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The Armv8-R AEM FVP entry path has carried TX_R52_BOOT_AT_EL1 since it was written, for the case its own comment describes: "an earlier boot stage or a vendor EL2 monitor has already dropped privilege to EL1". This board's entry path did not, so a kernel could not be built as a guest on it at all -- and it is the board where that matters most, because it is the one with silicon behind it. The bracket is the whole change. Everything from the Thumb reset trampoline to the ERET goes inside #ifndef TX_R52_BOOT_AT_EL1, and the #else supplies a one-instruction A32 _start that branches to el1_entry. A32 AND NOT T32, which is the one real difference from the standalone entry. The core resets in Thumb state here because the RTU boot instruction NXP plants is a T32 branch, but a guest is not reached by reset: it is reached by the monitor's ERET, and the monitor chooses the state through SPSR.T. Get the two out of agreement and the guest dies on its first instruction with an undefined-instruction exception, which looks exactly like a bad entry address and sends the reader to the loader instead of to the ERET. WHAT THE MONITOR INHERITS is enumerated at the #ifndef, next to the code it replaces rather than in a document, because that is where somebody adding a third board will be looking. This board's EL2 block is considerably larger than the model's, and each item on the list is something a guest at EL1 provably cannot do rather than something it merely does not: CNTFRQ is writable only at the highest implemented exception level and reads zero out of reset; HCPTR.TCP10/TCP11 reset set, trapping every EL1 floating-point access; HSCTLR.TE is an EL2 register (SCTLR.TE is EL1's, and el1_entry still clears it below); ICC_HSRE.SRE makes every other ICC_* and ICH_* register exist at all; the low-latency peripheral port enables reset to zero and an EL1 write to that register traps to EL2; and the TCM enables are per-core with ENABLEEL2 SILENTLY IGNORED from EL1 -- measured on both BTCM and CTCM, the base took and bit 0 took while bit 1 stayed clear. CNTHCTL.PL1PCTEN and PL1PCEN are the deliberate omission from that list, and the note says why. This path opens both, because a standalone kernel owns the physical timer. A monitor that TIME-partitions its guests must not: a partition's physical time keeps running while it is descheduled, so a guest reading it can observe that it was not running. That is the monitor's decision rather than this file's, which is why the list says what a guest cannot do rather than what a monitor should. Verified both ways. The three standalone images build and link unchanged, and a kernel built with the option boots at EL1 on a S32Z280-594EVB under an EL2 monitor, runs two threads through a queue and a semaphore, and reports back -- with no other change to the kernel or to its port. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
57390fc0fe |
Refused the Cortex-R52 VFP option without a hard float ABI (#686)
TX_R52_ENABLE_VFP with the default soft float ABI cannot build. The option defines TX_ENABLE_VFP_SUPPORT, which enables the VMRS, VSTMDB and VLDMIA blocks in the port assembly, and -mfloat-abi=soft leaves the assembler with no FPU to accept them. The configuration fails with eight errors of the form tx_thread_system_return.S:118: Error: selected processor does not support `vmrs r4,FPSCR' in ARM mode none of which mentions the float ABI, so a user has to reason from VMRS back to the option that enabled it. The guard that exists said otherwise. It warned that "the compiler will not emit floating-point instructions, so the VFP context path will never be exercised", which describes a build that succeeds and is merely pointless -- and then let configure finish, so the warning scrolled past well before the assembler errors appeared. It is now a FATAL_ERROR that names the fix, which is what the same file already does six lines above for TX_R52_ENABLE_FIQ_NESTING without TX_R52_ENABLE_FIQ. That combination is rejected for being "meaningless", while this one, which cannot assemble at all, was only warned about. The severities were the wrong way round. The option's definition also moves below the check, so the block reads like the FIQ nesting one. The ABI is not promoted to hard automatically. TX_R52_FLOAT_ABI is a cache variable the user may have set deliberately, and silently overriding an explicit choice is worse than refusing a combination that cannot work. readme_threadx.txt carried the same claim, and its option list marked the FIQ nesting dependency inline but not this one. Both corrected. No regression test. Nothing in the tree asserts a configure-time failure -- there is no harness for it, and the sibling FIQ nesting guard has none either -- so a test for this would have to introduce that mechanism for one case. Verified by hand in both directions instead: the soft-ABI combination now stops at configure with the message above, and the hard-ABI feature build (VFP, FIQ, IRQ nesting, FIQ nesting) builds its eight images clean and passes ctest 8/8 on the Armv8-R AEM FVP. scripts/check_gcc.sh passes unchanged; it configures the Cortex-R52 CMake stage without TX_R52_ENABLE_VFP, so the new branch is not on its path. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
5b94bad6a2 |
Fixed the AArch64 samples, none of which had ever linked with GCC (#673)
Every AArch64 gnu example build failed at the sample link, all 27 of them --
13 under ports/ and 14 under ports_smp/:
libg.a(libc_a-init.o): in function `__libc_init_array':
undefined reference to `_init'
relocation truncated to fit: R_AARCH64_CALL26 against undefined
symbol `_init'
libg.a(libc_a-fini.o): in function `__libc_fini_array':
undefined reference to `_fini'
build_threadx_sample.sh links with -nostartfiles, which is correct for a port
carrying its own reset path, and that drops crti.o and crtn.o along with
everything else. startup.S calls __libc_init_array by design, and newlib's
implementation calls _init, which crti.o is what defines. The AArch32 scripts
are unaffected: they use nosys.specs and never reach __libc_init_array.
The fix links crti.o and crtn.o explicitly, bracketing the object list -- the
first must precede every .init contribution and the second must follow all of
them, so their position is load-bearing rather than stylistic. Both paths come
from the compiler's own -print-file-name, so nothing here hard-codes a
toolchain layout.
The atfe branch sets both to empty, deliberately: picolibc's __libc_init_array
does not call _init, those 27 images link today, and adding crti.o would change
a working link for no reason. That is also why check_clang.sh is green on these
and does not list them as expected to fail -- the LLVM path never reached the
gap, so nothing has ever linked them and failed.
Fixed in ports_arch/ARMv8-A/threadx/ports/gnu/example_build, which is the
single source for both the ports/ and ports_smp/ copies, then regenerated with
update.sh --port-sets tx,tx_smp. The 27 generated copies are in this commit
because ports_arch_check compares them.
Verified: all 27 link with arm-gnu-toolchain 14.3.rel1 aarch64-none-elf, where
0 of 27 did before; _init and _fini disassemble to the expected crti prologue
and crtn epilogue over a ret; check_clang.sh with ATfE 22.1.0 is still green on
all five stages, including the 42 script-driven example builds; check_ports.sh
is green including the reproducibility check.
No regression test: these are link-only example images that no host test
executes. What guards them is check_clang.sh's example stage today, and
check_gcc.sh's, which is the next change and is the reason this was found.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
ca62edd27d |
Swept the stack-heavy measurement across placements, and qualified its result (#638)
#636 reported that a stack in BTCM gave a threefold tighter spread than DRAM0 for stack-heavy work. That measurement used a single code placement, which is the methodology #631 and #633 exist to correct: the cache benchmark got an alignment sweep and the interrupt handler got one, and this measurement never did. It was noticed when #637 added two threads to the same image and the figure moved -- both spreads came out near 6500 and the minima rose 15%. The recursive body is now generated at four placements and all four are measured, per placement, in one image. placement BTCM min / spread DRAM0 min / spread offset 0 41854 / 6850 42036 / 6880 offset 16 48670 / 1946 48792 / 1978 offset 32 42388 / 6914 42752 / 7018 offset 48 48914 / 1860 49198 / 6786 Reproducible across runs to within a few hundred cycles. Spread is dominated by code placement rather than by the memory holding the stack. It ranges from 1860 to 6914 depending on where the body falls in a cache line, and placement also moves the minimum by 17%, from 41854 to 49214. Against that, the memory contributes a consistent but small advantage: BTCM's minimum is lower at all four placements, by 0.4% to 0.9%. BTCM's spread beats DRAM0's decisively at one placement of the four, offset 48, at 1860 against 6786. At the other three the two are within 2% of each other. So the effect #636 reported is real where it occurs and is not a property of the part: quoting it as one invited the reader to expect it everywhere. #636's claim should be read as qualified by this. A stack in BTCM buys a small consistent improvement in the best case and a large improvement in spread at some code placements and not others. Anyone building a determinism argument on it needs the placement sweep in the loop, not a single figure. The pad nops that displace each placement execute on every recursion level rather than once, so each placement carries a slightly different constant cost, about 0.6% at the widest. That cancels in the BTCM against DRAM0 comparison, which is made at the same placement, and does not affect spread within one. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
1ef49338e9 |
Gave each thread its own MPU window, and made a violation fault (#637)
First step towards a ThreadX module port for this core: establish that PMSAv8-R
regions can be switched per thread on this part, what that costs, and that a
violation actually faults. Those are the questions worth answering before
writing a module manager on top of them.
Two threads each own a 4 KB window at the top of DRAM2. The windows are carved
out of the broad data region in mpu.c, because isolation is only meaningful in
memory no other region already covers -- every other region in that map is a
wide RW window, so a private buffer inside one of them would be reachable by
every thread whatever else was programmed.
Each thread writes its own window, which must succeed, and then reaches for the
other thread's, which must fault. The second half is the part that matters: a
test that only shows a thread reaching its own memory would pass just as well
with no protection at all.
Measured on the S32Z280-594EVB, reproducible across three runs:
thread 0 window 0x3187E000 own: reachable other: faulted
thread 1 window 0x3187F000 own: reachable other: faulted
region switch cost: 562 to 604 cycles
The cost is worth noting for the module port to come. A context switch on this
part is about 1400 cycles, so switching one region adds roughly 40% to it, and
most of that is the dsb and isb rather than the register writes. A module switch
programming several regions should therefore batch the barriers once at the end
rather than per region.
Scope, stated plainly. The window is applied by the thread calling
thread_mpu_activate, not by the scheduler. The port's scheduler does call
_tx_execution_thread_enter under TX_ENABLE_EXECUTION_CHANGE_NOTIFY, which would
make it automatic, but that macro is read by port assembly compiled into the
shared threadx library, so enabling it would oblige all nine example targets in
this port to supply the four execution hooks. A ThreadX module port carries its
own copies of the port assembly for exactly that reason, and that is where the
switch belongs. There is no user mode, no syscall boundary and no loader here.
The fault is survivable the same way the boot probes make it survivable:
fault_expected tells the data abort handler to record the violation and resume
after the faulting access. That works in thread context because the handler
returns where it came from rather than to a fixed recovery point.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
f535a4ec67 |
Measured stack-heavy work against the memory holding the stack (#636)
#635 found BTCM worth about 7.4% on a context switch with no determinism advantage, and said why: a switch saves sixteen registers, roughly one cache line, so the stack's cache state has almost nothing to contribute. It named the interesting case as work with a large stack working set, said it had not been measured, and said it should not be assumed. This measures it, and the answer inverts the earlier one. deep_touch recurses 24 frames, writing a frame on the way down and reading it on the way up, so the working set is the whole descent. The cache is cleaned and invalidated before each sample, so every descent starts cold. Two threads, one stack in BTCM and one in DRAM0, no partner threads and no relinquish: the timed region is entirely within one thread. Reproducible across runs: min mean max spread stack in BTCM 42868 43031 44890 2020 stack in DRAM0 42982 43291 49842 6860 The mean is the same to within 0.6%. The worst case is 10% lower for BTCM and the spread is 3.4 times tighter. No sample in either configuration exceeded twice the minimum, so these maxima are the workload rather than a timer tick -- which is the mistake that produced a false jitter result in #635 and is why the count of interrupted samples is printed. Put beside #635 the two measurements say opposite things and both are true. For a context switch, a small footprint touched every time, BTCM buys throughput and no determinism. For stack-heavy work, a large footprint touched once, it buys determinism and almost no throughput. The reason is that this workload is compute bound at the optimisation level this BSP builds at: 43000 cycles for 24 frames is dominated by call and loop overhead, so line fills are a few percent of the total and barely move the mean. What they do is vary, and that variance is what a bank with no cache in the path removes. So the determinism argument for TCM holds here, but it is worth 10% of worst case and a threefold narrowing of spread, not an order of magnitude. Anyone citing this in a safety argument should cite those numbers and not a larger claim. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
367d91880b |
Measured context-switch cost against the memory holding the stack (#635)
#634 placed a thread stack in BTCM and deliberately claimed no timing benefit, because none had been measured. This measures it. Two pairs of equal-priority threads hand control back and forth with tx_thread_relinquish. One pair has both stacks in BTCM, the other in DRAM0, and the measuring thread of each pair times the round trip in PMU cycles. Both pairs run in one image from one copy of the measuring code, which is what makes the comparison safe: the alignment trap that invalidated earlier work here bites when two builds with different layouts are compared, and a code shift moves both pairs equally. Reproducible to the cycle across runs: min mean max BTCM stacks 1370 1379 1402 DRAM0 stacks 1476 1480 1508 BTCM is about 7.4% faster, or 110 cycles on a round trip of two switches. Three findings that bound the claim, and the last one deflates it. The figure holds whether the cache is warm or cold. Cleaning and invalidating the data cache before every timed switch costs both configurations about 40 cycles and leaves the gap at 7.4%: warm it is 1334 against 1440, cold 1370 against 1476. So the advantage comes from BTCM's zero wait states, not from avoiding cache misses. That is because a context switch touches almost no stack -- sixteen registers, about one cache line -- so the stack's cache state has little to contribute either way. TCM should matter much more for threads with deep call chains or large locals, where the stack working set is big enough for cache state to dominate. That is not measured here and should not be assumed. There is no determinism benefit visible in this test. Excluding preempted samples, jitter is 32 cycles for BTCM and 28 to 34 for DRAM0 -- comparable, not better. A first version of this measurement appeared to show BTCM with 16 times less jitter, and that was wrong: max was reporting whichever pair a timer tick had landed on. Across three runs the outlier appeared in the BTCM pair once and the DRAM0 pair twice. Samples past 2000 cycles are now counted separately and excluded from min, mean and max alike, and the count is printed so the reader can see how many there were. Also tried and discarded: loading the partner thread with a cache walk to create pressure. The timed round trip includes the partner, so the walk dominated every sample and put all 256 past the outlier threshold. The per-sample flush replaced it and sits outside the timestamps. The demo also starts the PMU cycle counter, which bsp_boot.c does for the probe image and this image never ran. Without it every reading would have been zero, which reads as a free context switch rather than as a dead counter. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
62e966f6bf |
Enabled BTCM, measured it, and put a ThreadX thread stack in it (#634)
* Enabled BTCM and measured what it offers as a data store
BTCM was disabled because of a regression that turned out not to exist; the
claim was retracted in the previous commit. It is enabled now, and this
measures why that is worth doing: 16 KB at zero wait states, where ATCM has
one, and no cache in the path at all.
Enabling it needs three things, all of which existed for ATCM already: the
region register write at EL2, an MPU region, and an ECC preload before any read
(TRM 6.2.2). BTCM accepts 32-bit stores where ATCM needs 64-bit, which
tcm_preload already handles.
The cache benchmark now sweeps three memories rather than one, four loop
alignments each. At the alignments where the loop is not instruction-fetch
bound:
memory cold (uncached) warm (cached) gain
DRAM2 half-speed 857,540 651,436 24.0%
DRAM0 full-speed 797,824 651,297 18.3%
BTCM zero wait 694,689 651,369 6.2%
Three things follow.
Warm times are identical across all three memories, within 0.02%. Once the data
cache is working the backing store barely matters, because the working set fits
in it.
Cold times rank as the reference manual predicts: BTCM fastest, then DRAM0,
then DRAM2 at half the core frequency (S32Z2 RM 6.3.6).
BTCM still shows a 6.2% gain when the caches are enabled, and that cannot be
the data cache, because an enabled TCM is Non-cacheable Non-shareable Normal
memory whatever the MPU says. It is the instruction cache on the timing loop.
This probe has always measured both caches together; three memories side by
side is what makes that visible.
The number that matters for placing data in BTCM: uncached BTCM is within 6.6%
of the best cached case, where uncached DRAM0 is 22% off it. Data in BTCM runs
at close to cache-hit speed with no cache to miss, which is the determinism
argument stated as a measurement rather than an assertion.
DRAM2's sweep is unchanged with BTCM enabled, 0 and 0 and 240 and 240, which
independently confirms the retraction.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Put a ThreadX thread stack in BTCM
The code side of TCM was done in #630; this is the data side. One of the demo's
three thread stacks now lives in BTCM and the other two stay in DRAM0, so a run
exercises both paths and a mistake in either shows up.
link.lds gains a BTCM region and a .btcm_bss NOLOAD section, so a stack is an
ordinary C array with a section attribute and the linker checks it fits, rather
than a hardcoded address that silently overflows the bank.
entry.S preloads the whole bank at EL2, and that is not optional. ECC is enabled
on this part, so a TCM location must be written before it can be read (TRM
6.2.2), and a stack is read before the program writes it -- the first context
restore pops what tx_thread_create built into it. The preload has to happen
before any C runs, because the demo images do not run bsp_boot.c, which is where
the ATCM preload lives. 32-bit stores suffice for BTCM where ATCM needs 64-bit.
Why BTCM for a stack: 16 KB at zero wait states where ATCM has one, and never
cached whatever the MPU says about it. Measured in the previous commit, uncached
BTCM comes within 6.6% of the best cached case while uncached DRAM0 is 22% off
it, so stack access runs at close to cache-hit speed without depending on a line
being resident. That is the property a determinism argument needs.
What this commit does not claim: no thread-level timing improvement has been
measured. The case for BTCM here rests on the memory characterisation and on
removing the cache from the path, not on a measured context-switch figure. That
measurement is worth doing and has not been done.
Verified on the S32Z280-594EVB. The demo reports its stack addresses so the
placement is visible rather than implied -- sleeper at 0x30100000 in BTCM,
spinner and judge in DRAM0 -- and passes with 100 ticks, 20 sleeper wakeups and
20 preemptions, so a real thread schedules, preempts and context-switches on a
tightly-coupled-memory stack. The boot image still passes six of six probes.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
35de56853c |
Retracted the claim that a second TCM bank costs the data cache (#633)
entry.S and readme_s32z280.txt both stated that enabling any second TCM bank removes all measurable data-cache benefit on this part, and gave five configurations as evidence: ATCM alone at a 24% cache gain, four combinations involving a second bank at none. Both concluded the cause was not a particular bank, not its address and not the ECC preload, but enabling a second bank at all. Both cited the Cortex-R52 and S32Z2 errata as not covering it and offered a shared RAM pool between the LLC and the TCMs as an explanation. A defect report went to NXP on that basis. It was an artifact of the benchmark. That benchmark was bimodal with respect to where its timing loop fell inside a 64-byte cache line, reporting either 24% or nothing at all for identical silicon, and every one of those five configurations was an edit to entry.S, so every one shifted the code that followed and moved the loop between modes. Adding two nop instructions reproduces the "second bank" figure exactly, to the digit. The report to NXP has been withdrawn. Re-measured with the alignment sweep added in #631, one bank and two are indistinguishable: loop offset in line ATCM only ATCM + CTCM 0 gain 0 gain 0 16 gain 0 gain 0 32 gain 240/1000 gain 240/1000 48 gain 240/1000 gain 240/1000 So enabling a second bank costs nothing measurable. The banks stay disabled, but for the ordinary reason that nothing in this example uses them, and both texts now say that instead. Enabling one is a single line, with the ECC preload before any read (TRM 6.2.2) and an MPU region as the only prerequisites, both already handled for ATCM. The readme also now states the general point, which outlasts the TCM detail: a single-figure timing result from this example cannot be compared across builds unless the timed loop's alignment is controlled, because almost any change shifts code. Comments and documentation only; no generated code changes. Verified on the board regardless, since entry.S was touched: six of six probes pass and the sweep is unchanged. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
3aaaa6700d |
Revalidated the ATCM handler result across four placements, and it holds (#632)
The handler comparison in #630 measured one alignment, which is the mistake that made the cache benchmark in this example report 24% or 0% for identical silicon. The handler body is now generated at four offsets within a cache line, all four are measured, and the figures are reported per placement. The claim survives. Mean cycles for the handler body: loop offset code RAM ATCM 0 523 383 16 534 377 32 539 375 48 521 383 ATCM is faster at every placement, by about 28%, and the two sets of means do not overlap. Worst case improves as well, 510 against 694. Two things worth recording beyond the headline. The handler measurement is only mildly alignment sensitive, 3.5% across placements in code RAM and 2% in ATCM, quite unlike the cache loop's two modes. So this comparison was less fragile than the cache one, and #630's direction was right even though its method was not defensible. The absolute numbers differ from #630 because the body now sits behind a placement wrapper that adds a call; the comparison is internally consistent either way. Both variants also report identical cache sweeps, 0 and 0 and 240 and 240, which settles the regression this branch's predecessor appeared to show. That apparent regression was the single-alignment probe moving between its two modes, not anything about ATCM. One copy of the logic is kept: the wrappers inline a single always_inline implementation, so the four placements cannot drift apart. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
1d3f3f8e4c |
Measured the cache benchmark at four alignments, because one is not enough (#631)
This benchmark was bimodal and reported a single number, which made it worse
than no benchmark. The same workload on the same silicon reports either 24%
cache benefit or none at all, decided only by where the loop falls inside a
64-byte line -- and therefore by any unrelated change that shifts code
ahead of it. Two nop instructions added to entry.S were enough to flip it.
That is not a hypothetical. A run of conclusions drawn from this probe turned
out to be measuring code layout: an interrupt-handler comparison, a claim that
enabling a second TCM bank costs all cache benefit, and a follow-up claim that
what mattered was when the TCM region register was written rather than what it
contained. The last of those was reported to NXP as a defect and has had to be
withdrawn. Enabling CTCM and adding two nops produce identical results, to the
digit, because the only measurable consequence of the enable was the eight
bytes of instructions it added.
Four copies of the loop are now generated at different offsets within a cache
line, all four are measured, and the low and high gains are both reported.
Pinning a single alignment was tried first and is not a fix: it silently picks
one of the two modes -- aligned to 64 the loop sits permanently in the low one.
Measured on the S32Z280-594EVB, reproducing exactly across runs:
loop offset in line cold warm gain
0 890,302 890,035 0%
16 890,208 889,976 0%
32 857,439 651,390 24.0%
48 857,631 651,571 24.0%
The cold pass differs between the modes as well, 890k against 857k, so the
loop is slower even with both caches off. The cold pass is instruction-fetch
bound out of code RAM at half the core frequency (S32Z2 RM 6.3.6), and how the
loop straddles lines decides how much of the data cache's contribution is
visible at all. This probe therefore measures both caches together and always
did; the sweep at least makes the variation visible instead of letting one
arbitrary placement stand in for the part.
C4 now passes if any alignment shows a 10% speedup, and says so explicitly
when the low mode does not, so the sensitivity appears in the log rather than
being discovered later.
Verified: with the sweep in place, adding 0, 8, 12 or 20 bytes of nops to
entry.S leaves the reported low and high gains unchanged. Before it, the same
shifts read 24.0%, 0%, 0% and 0%.
Also adds cache_disable_all, which the sweep needs: cache_enable was one-way,
so a second cold reading in one run was impossible.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
b2983c08d3 |
Ran the interrupt handler from ATCM, and measured what that buys (#630)
The TCM work so far enabled ATCM and left it empty, which buys nothing.
This places code in it and measures the result.
link.lds gains an ATCM region and an .atcm_text section whose run address
is in the bank and whose load address is in CODE. tcm_copy_atcm_text
moves it, using 64-bit stores because ECC is enabled on this part and
ATCM requires them (Cortex-R52 TRM 6.2.2); both ends of the section are
8-byte aligned so there is no narrower tail to leave without check bits.
The copy runs after T4 and T5, which write test patterns to the first and
last words of the bank and would otherwise land on top of the code.
s32z280_atcm.elf is the same image as s32z280_boot.elf with the interrupt
service body placed in ATCM. Both targets exist so the comparison can be
repeated on one board in one session without reconfiguring. The service
routine is split into a timed wrapper that stays in .text and a body that
moves, so the wrapper's own cost appears in both measurements and cancels.
Measured in PMU cycles over 64 samples, caches enabled in both:
code RAM ATCM change
min 454 334 -26.4%
mean 458 340 -25.8%
max 612 466 -23.9%
spread 158 132 -16.5%
CNTPCT is not used for this: at 8 MHz it cannot resolve a handler body,
let alone the variation in one.
The level shift is the solid part. ATCM is a quarter faster even though
the caches were on and code RAM had the instruction cache available,
which says the handler does not stay resident between interrupts 10 ms
apart -- so each one pays a cold fetch from code RAM, which runs at half
the core frequency where ATCM runs at full speed with one wait state
(S32Z2 RM 6.3.6).
The determinism claim deserves less weight than the numbers first
suggest. The spread narrows by only 16%, and ATCM's worst case still sits
slightly above code RAM's best case, so the two distributions overlap at
the tails rather than separating. Whatever jitter remains is not
dominated by instruction fetch.
Both images pass six of six boot probes.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
ff638dae71 |
Configured the LINFlexD console twice, because once is not enough at -O2 (#629)
The console driver runs its configuration sequence a single time, and that works only because this BSP is built without an optimisation flag. Compiled at -O2 the same sequence leaves the line corrupted: every character partially wrong, in the pattern the file already describes for a misconfigured module. What makes it worth guarding against is that the failure is invisible. UARTCR reads back exactly the value written. LINIBRR and LINFBRR read back exactly the values written. linflexd_init returns LINFLEXD_INIT_OK. The registers are right and the line is wrong, so nothing in the returned status tells the caller the console cannot be trusted. Localised by bisection: with every other file at -O2 and this one at -O0 the output is clean, and with only linflexd_init at -O2 it is corrupted, so the fault is in the configuration sequence rather than in the per-byte transmit path. The mechanism is not understood, and this commit does not claim to explain it. Tested and rejected: a 100x larger bound on the wait for initialisation mode, a settling delay before the first LINSR read, a settling delay after leaving initialisation mode, a barrier and read-back between the two UARTCR writes, and waiting for LINSR to report the exit from initialisation mode. None of those makes a single pass work at -O2. A second pass does, at both optimisation levels, which is what this does. Instrumented with a duplicate of the sequence forced to -O2 and reported through a console repaired afterwards, which is how the register read-backs above were obtained. Verified on the S32Z280-594EVB. The boot image passes six of six probes with the console status still reporting 0x00000000, and the reproducer builds clean and prints correctly at both -O0 and -O2. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
0e7db2bf79 |
Stopped the generated Armv8-M readmes claiming a false origin date (#623)
Every Armv8-M port readme ends with
09-30-2020 Initial ThreadX 6.1 version for Cortex-M85 using GNU tools.
with the core name substituted in by scripts/copy_armv8_m.sh. The date is the
shared port's, so each core inherits it whatever its own history: Cortex-M85 was
announced in 2022 and its readme claims a 2020 origin, and any core added later
gets the same treatment the moment its name joins the generator's list.
The rest of the history block is accurate, since it records changes to the shared
files. Only the closing line asserts something per-core. Reword it to describe
the Armv8-M port itself, and say where a given core's real starting point is.
Regenerating updates the twelve readmes for cortex_m33, cortex_m52, cortex_m55
and cortex_m85 across the three toolchains.
Cortex-M52 makes the point: it arrived in #519 and its readme immediately claimed
a 2020 origin for a core announced in 2023.
The ARMv7-M templates say "Initial ThreadX version 6.1.7 for Cortex-M", with no
placeholder to substitute, so they make no per-core claim and are left alone.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
090f8edc56 |
Added Cortex-M52 (Armv8.1-M) port (#519)
* ports: add Cortex-M52 (Armv8.1-M) port support Add ThreadX port for Cortex-M52, supporting three toolchains: - GNU (GCC) - AC6 (Arm Compiler 6) - IAR Cortex-M52 is an Armv8.1-M Mainline processor sharing the same architecture profile as Cortex-M55 and Cortex-M85. The port is functionally identical to the existing Cortex-M85 port. * ports: update cortex_m52 port copy script,using it update content -Add cortex_m52 in threadx/scriprts/copy_armv8m.sh -Using updated copy_armv8m.sh to generate new cortex_m52 port content * Regenerated the Cortex-M52 port against dev and wired it into the checks The port was generated from ports_arch/ARMv8-M as it stood on master, which has since moved on. Rebasing onto dev and running scripts/copy_armv8_m.sh again brings the twelve stale files into line, which is the point of generating them: the core picks up every ARMv8-M fix made since without anyone porting it by hand. Among what it picks up: "MOV r0, 0" becomes "MOV r0, #0" in the schedule and system-return paths, the non-canonical immediate form that GNU as tolerates and LLVM's assembler rejects; and gnu/src/tx_initialize_low_level.S goes away, since the shared source no longer has it. Two integration points exist only on dev, so the original change could not have included them. cmake/cortex_m52.cmake, so the port can be selected the documented way. Every other Cortex-M core has one. It uses the hard float ABI, as Cortex-M55 and Cortex-M85 do. An entry in scripts/check_clang.sh, likewise with -mfloat-abi=hard. That flag is not decoration: -mcpu=cortex-m52 implies Helium, and building it soft-float ends in "multilib configuration error: No library available for MVE with soft-float ABI" on every file, which reads as a broken port rather than a missing flag. Verified after regenerating: scripts/copy_armv8_m.sh is a no-op, so the tree matches its source; all 14 assembly sources and 188 C sources compile for cortex-m52 with Arm Toolchain for Embedded 22.1.0. Note for anyone building with GNU tools: arm-none-eabi-gcc 13.2.1 rejects -mcpu=cortex-m52 outright. Support arrives in GCC 14. --------- Co-authored-by: Frédéric Desbiens <frederic.desbiens@eclipse-foundation.org> Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
09c190a71d |
Added a CMake target for the Linux sample program (#622)
The CMake build produced libthreadx.a and nothing else, so trying ThreadX on
Linux meant using the Makefile beside the port instead. Build the demo the
Makefile builds, for the linux port and its SMP counterpart.
The target is behind an option that defaults off, so an ordinary build is
unchanged and still produces just the library. -DTHREADX_SAMPLE=ON adds it:
cmake -S . -B build -DTHREADX_ARCH=linux -DTHREADX_TOOLCHAIN=gnu \
-DTHREADX_SAMPLE=ON
cmake --build build --target sample_threadx
The include path uses TX_COMMON_DIR rather than naming common or common_smp,
since the top level already resolves which of the two applies.
Verified by building and running both variants. Non-SMP prints
**** ThreadX Linux Demonstration **** (c) 1996-2020 Microsoft Corporation
and SMP prints the SMP banner, both with the demo's thread counters advancing. A
default configure with no THREADX_SAMPLE has no sample_threadx target and still
produces libthreadx.a, so nothing existing moves.
Derived from the two example_build files in #404 by Yanfeng Liu, which had the
same goal. That change also rewrote the top level's SMP selection, added
common_smp/CMakeLists.txt and added ports_smp/linux/gnu/CMakeLists.txt; all three
have since arrived on dev by other routes, so only the sample targets were still
missing. The include path needed adjusting because the original depended on
THREADX_SMP being a string suffix, which it no longer is.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
1296cf1740 |
Corrected the S32Z280 data SRAM map against the Reference Manual and the part (#621)
The RTU has 1 MB of data SRAM in three contiguous banks, and this port described it wrongly in both directions. DRAM0 0x31780000 256 KB full core speed DRAM1 0x317C0000 256 KB full core speed DRAM2 0x31800000 512 KB half core speed DRAM1 was not declared at all, so 256 KB of full-speed memory went unused and the linker's DATA region stopped at 256 KB. DRAM2 was declared as 2 MB when it is 512 KB, which mattered more: the MPU mapped 1.5 MB past the end of the bank, and that range aliases back onto its base. Anything placed above 0x31880000 would have shared storage with the bottom of the region silently -- no fault, two objects at one address. Nothing was placed there yet, so this was a trap rather than a live defect. Sources: S32Z2 Reference Manual Rev. 5, section 6.3.6 and Table 13, and the board. Writing distinct values to all three banks and reading them back shows 1 MB of independent storage, and the first word past DRAM2 returns the value written to its base, which is what fixes the size. Note that NXP's own debugger memory map, s32z2e2_memory_regions.py in S32 Design Studio, calls the last bank 2 MB. The Reference Manual and the silicon agree it is 512 KB. Also corrected the description of these banks throughout. They are all RTU-local; the earlier comments treated locality as the thing that distinguishes them, when the actual difference is clock speed. That is why the cache benchmark uses DRAM2 -- caching a bank that already runs at core speed shows nothing, which is a real effect the old wording explained with the wrong cause. Verified on the S32Z280-594EVB: builds clean, and the boot probes pass six of six with the MPU, GIC, interrupts, caches and both protection faults exercised. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
358e7a9ea5 |
Stopped the remaining example scripts naming archives that were deleted (#620)
#618 fixed the arm9 and arm11 shell scripts, but not their .bat counterparts, and did not look at cortex_r4 and cortex_r5 at all. An audit of every build script against the files actually present in its directory found the rest. The .bat scripts for arm9, arm11, cortex_r4 and cortex_r5 still linked libc.a, libgcc.a and, for arm11, libnosys.a from their own example_build directories, and the cortex_r4 and cortex_r5 shell scripts did too. Those archives went in 6.1.10 under "Removal of unneeded files", so on Windows all four examples failed exactly as the shell versions did before #618, and on Linux the two R-profile ones still did. Link through the compiler driver, as the other examples have since #594. This does not make the examples link, and the change stops there deliberately. All four now fail the same way, undefined reference to `_fini' because their linker scripts define the .init and .fini sections but not the _init and _fini symbols, which live in crti.o and crtn.o and are omitted by -nostartfiles. Reviving four very old cores is separate work. Correct the comment on EXAMPLES_EXPECTED_TO_FAIL again. It had cortex_r4 and cortex_r5 failing for want of newlib multilib variants; they do not. All four share the single cause above, and the multilib explanation was wrong for the R-profile pair just as it was for arm9 and arm11. Verified by running each shell script: cortex_r4 and cortex_r5 fail on _fini rather than on missing files, matching arm9 and arm11. The .bat changes mirror link lines proven that way in the same directories; they cannot be run here. Three scripts are left alone and reported instead, because a blind edit could not be verified: ports_module/cortex_m3 and cortex_m4 have Windows-only .bat scripts that also name sources which are absent or differ in case, and ports_smp/mips32_interaptiv_smp needs a MIPS toolchain. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
20d4b977f2 |
Removed generated build output and per-user IDE state, and fixed two stale link lines (#618)
* Removed generated build output and stopped two scripts naming deleted archives Three kinds of file in the tree are produced by a build rather than written by hand, and one pair of scripts still links archives that were deleted years ago. Keil writes ThreadX_Library.plg on every build; the two committed copies are HTML build logs from someone's machine. Code Composer generates the makefiles under ports/c667x/ccs/example_build/tx/Release from the project files beside them, so makefile, objects.mk, sources.mk, subdir_rules.mk, subdir_vars.mk and ccsObjs.opt are all regenerated output. The arm9 and arm11 sample builds link libc.a, libgcc.a and, for arm11, libnosys.a from their own example_build directories. Those archives were removed in 6.1.10 under "Removal of unneeded files", and libnosys.a in #594, but the link lines were never updated, so both examples fail immediately with arm-none-eabi-ld: cannot find libc.a: No such file or directory Link through the compiler driver instead, the shape every other example in the tree uses since #594: the driver supplies libc and libgcc, and SYSCALL_LIB is already defined in both scripts. That does not make either example link, and the fix stops short of that on purpose. With the archives no longer named, both now fail on undefined reference to `_fini' because their linker scripts define the .init and .fini sections but not the _init and _fini symbols, which live in crti.o and crtn.o and are omitted by -nostartfiles. Making those two old cores build is a separate question from removing a stale reference, so they stay in EXAMPLES_EXPECTED_TO_FAIL, with the comment there corrected: it blamed newlib multilib packaging, which is true of cortex_r4 and cortex_r5 but was never the reason for arm9 and arm11. Reproduced throughout with arm-none-eabi-gcc 13.2.1. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> * Removed the Keil per-user state files and ignored them Every Keil project in the tree carried a second file holding per-user state: 25 .uvoptx beside the 25 .uvprojx, 9 .uvopt beside the 9 .uvproj, and 3 .uvgui multi-project workspace files. uVision rewrites all of them whenever a project is opened, so they record whoever last had it open rather than anything about the port: debugger selection, breakpoints, watch windows, window geometry. Nothing in the tree references them, and every affected directory keeps its .uvprojx or .uvproj, which is the file that actually describes the project. Add ignore rules so they do not come back the next time someone opens a project and commits. 1.4 MB across 37 files. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
b1a48824ea |
Enabled ATCM on S32Z280, and left the other two banks off for a measured reason (#617)
entry.S programs ATCM to 0x30000000 and enables it at both exception levels. This
happens at EL2 deliberately: writing ENABLEEL2 from EL1 is silently ignored, which
was measured on BTCM and CTCM -- the base took and ENABLEEL10 took while
ENABLEEL2 stayed clear -- and the same write from EL2 sticks. tcm_enable() keeps
ENABLEEL10 as its success criterion for the same reason, since requiring both
would report failure for a bank that is usable at the level the caller runs at.
Programming ATCM also moves it off address 0, where CFGTCMBOOTx leaves it, so a
null-pointer write now faults instead of quietly landing in tightly-coupled
memory.
The boot image verifies rather than programs, preloads ATCM because ECC is enabled
and the check bits are not initialised by the core, and proves the bank holds data
both before the MPU is enabled and after. One MPU region covers it: the TRM
requires a region before an enabled TCM can be used, and an enabled TCM always
behaves as Non-cacheable Non-shareable Normal memory whatever the region says, so
only the permissions there matter.
BTCM and CTCM are left disabled, and that is a measurement rather than caution.
Enabling any second bank removes all measurable data-cache benefit:
ATCM only cache gain 24%
ATCM + BTCM cache gain 0%
ATCM + BTCM + CTCM cache gain 0%
ATCM + BTCM at another base cache gain 0%
ATCM + CTCM, BTCM disabled cache gain 0%
Five configurations, one variable. Not a particular bank, not its address, and not
the ECC preload: enabling a second bank at all. The benchmark buffer is in
non-RTU-local SRAM at 0x31800000, outside every TCM window, and CCSIDR reports the
same 16KB four-way cache throughout. I was wrong twice while narrowing this --
first blaming the preload, then blaming BTCM specifically -- and each was settled
by a run rather than by argument.
No erratum covers it. Checked the Cortex-R52 errata notice SDEN-857344 issue 19,
all twenty-five entries, and the S32Z2 0P91J mask set errata, whose RTU and R52
entries are ERR050509, ERR051107, ERR051153, ERR051441, ERR051613, ERR051614 and
ERR052126. A RAM pool shared between the RTU's last-level cache and the TCMs would
explain it, the LLC being documented as allocating ways to specific domains, but
that is a guess and it belongs with the other questions for NXP.
Little is lost meanwhile. The reference manual describes TCM_A as the bank
"optimized for small, regularly executed code such as interrupt service routines
or OS kernels", which is what a TCM is wanted for here, and enabling the other two
is one line each in entry.S once there is an answer.
Two checks were also wrong and are fixed. T5 reported every bank accessible while
two were disabled, because a disabled TCM's address range is serviced through AXIM
and memory answering there says nothing about the TCM; it now requires the bank to
be enabled as well. And T2 called tcm_enable() from EL1 for all three banks, which
would have switched on the very banks entry.S leaves off.
Verified on S32Z280 silicon: ATCM enabled and holding data before and after the
MPU, the cache benchmark back to 24%, protection checks X2 and X4 unchanged, and
the ThreadX demo still reporting 100 ticks with 20 preemptions.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
27891f3b84 |
Read the TCM configuration off the part, and corrected what we thought we knew (#615)
The BSP has recorded since bring-up that the TCMs are inaccessible at reset
"because nothing has programmed the TCM region registers yet". Reading those
registers shows the reason was wrong.
ATCM 0x0000011F 64KB, 1 wait state, ENABLED at EL2 and EL1/EL0
BTCM 0x00000014 16KB, 0 wait states, disabled
CTCM 0x00000114 16KB, 1 wait state, disabled
Every BASEADDRESS field is zero, so ATCM is live as 64KB at address 0x00000000,
not at the 0x30000000 the reference manual documents. The reads that faulted were
of an address the TCM is not at. ATCM's enables reset set because CFGTCMBOOTx is
tied high on this part, which the Cortex-R52 TRM gives as the one exception to
"at reset all bits are 0 apart from SIZE and WAITSTATES". BTCM and CTCM really
are disabled.
The sizes and wait states match the S32Z2 reference manual exactly -- TCMA 64KB
with one wait state, TCMB 16KB with none, TCMC 16KB with one -- so the TRM's field
layout, NXP's documented configuration and the silicon all agree. That agreement
is the point of reading before writing.
ECC is implemented and enabled: IMP_MEMPROTCTLR reads 0x00000011, both RAMPROTIMP
and RAMPROTEN set. TRM 6.2.2 therefore applies rather than being hypothetical: a
TCM location must be written before it is read, or the read reports an error --
which looks exactly like "the TCM is not accessible" and sends the reader back to
region registers that were already correct. The preload widths differ too, ATCM
needing 64-bit aligned STRD or STM where BTCM and CTCM accept 32-bit stores, so a
C loop over unsigned int would leave ATCM's check bits invalid.
tcm.c reads and decodes only; nothing is programmed here. The layouts in tcm.h are
quoted from TRM r1p3 section 3.3.94 table 3-136 and section 3.3.76 table 3-114,
not inferred from a neighbouring register: BASEADDRESS is [31:13] where
IMP_PERIPHPREGIONR uses [31:12], and assuming the analogy would have been wrong by
one bit in the same way the PRBAR shift was.
The boot image reports all of it and flags any size that disagrees with the
reference manual, so a part configured differently says so rather than being
silently assumed to match this one.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
ea91beae42 |
Verified nested FIQ handling on S32Z280 silicon (#614)
#613 exercised the FIQ nesting routines on the FVP. What the model could not show is whether a real GIC-600 routes Group 0 to FIQ the same way, which is the reason to run it here. It does, and the counts match the model exactly. The Group 0 support ports across unchanged: IGRPEN0 and BPR0 on the CPU interface, Group 0 in the distributor, gicv3_enable_sgi_group0, gicv3_send_sgi_group0 through ICC_SGI0R, and the separate Group 0 acknowledge and end-of-interrupt pair. entry.S routes the EL1 FIQ vector into _tx_thread_fiq_context_save with the acknowledge before nesting starts, and leaves FIQ unmasked on the drop to EL1 when FIQ support is compiled in, for the same reason as on the FVP: tx_thread_stack_build only clears a thread's F bit in that configuration. One structural difference from the FVP cost a link. This example reports faults through FAULT_TAIL rather than FAULT_REPORT, so the vector table needed a new el1_fiq_entry label that falls back to fault_el1_fiq. Placing that label inside the TX_R52_USE_THREADX_IRQ guard broke s32z280_boot.elf, which does not define it: the vector reference is unconditional, so the label has to be too. It now sits outside the guard and carries its own, the same shape the demo_m2 link break in #613 forced on the FVP side. Verified on S32Z280 silicon: F1 FIQ delivered and dispatched PASS low-priority FIQ count = 0x00000015 21 high-priority FIQ count = 0x00000014 20 nested FIQ count = 0x00000014 20 of 20 nested max FIQ depth = 0x00000002 FIQ depth now = 0x00000000 F2 FIQ nested inside an FIQ handler PASS F3 FIQ nesting unwound to depth zero PASS F4 IRQ tick undisturbed by FIQ work PASS F5 lower-priority thread still scheduled PASS F6 no unexpected Group 0 INTID PASS No regression, both configurations checked on the board. In the FIQ build the ThreadX demo still reports 100 ticks with 20 preemptions and the boot image still passes its cache and protection checks. In the default build the demo is unchanged and the image links no FIQ or Group 0 symbol at all, so the work is absent rather than dormant where it is not wanted -- worth confirming on hardware rather than reasoning about, because the default build now reaches the FIQ vector through a new label even though that label only branches to the fault reporter. entry.S assembles in all four combinations of TX_R52_USE_THREADX_IRQ and FIQ support, on both toolchains, and every file builds with GNU without warnings and with Arm Toolchain for Embedded 22.1.0. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
19d90a49c4 |
Exercised the nested FIQ path, the last pair nothing had ever called (#613)
_tx_thread_fiq_nesting_start and _tx_thread_fiq_nesting_end complete the set: after #611 and #612 covered IRQ nesting on the model and on silicon, these two were the remaining routines compiled into every build and entered by nothing. demo_fiq.elf enters them. FIQ needs more of the GIC than IRQ does. With a single security state the controller delivers Group 0 as FIQ and Group 1 as IRQ, so an interrupt only arrives as an FIQ if it has been moved into Group 0, the distributor and the CPU interface both have Group 0 enabled, and it is acknowledged through the Group 0 registers. Group 1's acknowledge returns the spurious INTID for a Group 0 interrupt and leaves it pending, which would present as a storm rather than as an error. gicv3.c gains IGRPEN0, BPR0, gicv3_enable_sgi_group0, gicv3_send_sgi_group0 and the Group 0 acknowledge and EOI pair. ICC_SGI0R differs from ICC_SGI1R only in opc1, 2 against 0, and each raises into its own group. entry.S routes the EL1 FIQ vector into _tx_thread_fiq_context_save with the same ordering the IRQ path needed: acknowledge in FIQ mode before nesting starts, then nesting_start, service, nesting_end, and end-of-interrupt last. It also leaves FIQ unmasked on the drop to EL1 when FIQ support is compiled in, because tx_thread_stack_build only clears a thread's F bit in that configuration and nothing else ever clears it, so an FIQ raised before the first thread ran would otherwise be silently ignored. Nesting an FIQ means taking an FIQ while an FIQ handler runs, which one source cannot show, so two Group 0 SGIs are used with the second at a numerically lower priority. The low one's handler raises the high one. One mistake worth recording, because the guard it needed is not obvious. TX_ENABLE_FIQ_SUPPORT is PUBLIC on the threadx target, so it reaches every image as soon as the library is built with FIQ -- including images that link no interrupt controller at all. demo_m2 is one of those: no gicv3.c, no irq_dispatch.c. Referencing gicv3_acknowledge_group0 and board_fiq_service from the FIQ vector broke its link outright. The vector is now gated on TX_R52_USE_THREADX_IRQ as well, which is how the IRQ vector has always been gated, and images without the infrastructure keep the fault reporter. Verified on FVP_BaseR_AEMv8R. F1 covers the whole Group 0 chain on its own -- IGRPEN0, the ICC_SGI0R encoding, IAR0 and EOIR0, the EL1 vector and F being unmasked -- so a failure there points somewhere other than the nesting routines. The demo reports 21 low-priority FIQs, 20 high-priority, 20 of them nested, max depth 2, depth unwound to zero, the IRQ tick undisturbed, threads still scheduled and no unexpected Group 0 INTID. All six images pass in the FIQ configuration, and the default build links no FIQ or Group 0 symbol at all, so the work is absent rather than dormant where it is not wanted. Both toolchains build every file, GNU with no warnings and Arm Toolchain for Embedded 22.1.0 as well. Not done here: the S32Z280 example is untouched, so FIQ on silicon is a separate change, as IRQ nesting was. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
5a7f4f8f7c |
Verified nested IRQ handling on S32Z280 silicon (#612)
#611 exercised the nesting routines on the FVP and established what breaks them: the interrupt must be acknowledged before nesting starts, or the still-pending level-asserted timer is retaken the moment IRQ is enabled and recurses until the stacks are gone. entry.S here carries the same ordering for the same reason, and irq_dispatch.c splits the same way, board_irq_service taking an already-acknowledged INTID while board_irq_handler keeps its old shape as the non-nesting entry point. What the model could not answer is whether a real GIC-600 agrees, and two things could have differed. The first is the number of implemented priority bits. Equal priorities do not preempt and it is the low bits that vanish, so if the timer and the SGI collapse to one value after truncation then nesting cannot happen at all -- and the test would fail without saying why. gicv3_priority_bits discovers the count by writing 0xFF to a priority byte and reading back which bits stick, board_init records the two effective values, and check P1 requires the SGI to still outrank the timer. This silicon keeps five bits, the same as the FVP, so 0xA0 and 0x50 stay distinct; that is now measured and reported rather than assumed. The second is whether an SGI raised on real hardware is delivered at all. ICC_SGI1R is a 64-bit AArch32 CP15 register whose encoding does not transcribe from the AArch64 alias, so check N1 raises one from thread context and requires delivery before nesting is involved. It arrives. Verified on S32Z280 silicon: priority bits = 0x00000005 timer effective = 0x000000A0 sgi effective = 0x00000050 P1 SGI outranks timer after truncation PASS N1 SGI delivered and dispatched PASS max depth = 0x00000002 nested SGIs = 0x00000032 depth now = 0x00000000 N2 SGI nested inside another handler PASS N3 nesting unwound to depth zero PASS N4 tick still advancing after nesting PASS N5 lower-priority thread still scheduled PASS N6 no spurious or unexpected interrupts PASS Fifty nested SGIs across fifty ticks, one per tick, and depth never exceeded two. No regression, checked in both configurations on the board. In the nesting build the ThreadX demo still reports 100 ticks with 20 preemptions and the boot image still passes its cache and protection checks. In the default build the image links no nesting symbols at all and the demo is unchanged. Both toolchains compile every file, GNU with no warnings and Arm Toolchain for Embedded 22.1.0 as well. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
559460bed9 |
Exercised the nested IRQ path, which had never once been entered (#611)
_tx_thread_irq_nesting_start and _tx_thread_irq_nesting_end have shipped in this port since it was written, compiled into every build, and nothing had ever called either one. Not on the model, not on silicon, not in any demo. demo_nesting.elf enters them. Provoking nesting needs two sources with different priorities. The generic timer PPI was already there; the second is an SGI, which a core can raise on itself. gicv3.c gains gicv3_enable_sgi and gicv3_send_sgi for that. ICC_SGI1R is 64-bit, so in AArch32 it is an MCRR rather than an MCR, and the AArch64 name the Cortex-A72 example uses, S3_0_C12_C11_5, does not transcribe to the AArch32 CP15 space. The encoding here is confirmed by check N1 in the demo, which raises an SGI from thread context and requires it to be delivered and dispatched. The order of the pairing is the whole difficulty, and getting it wrong does not fail gracefully. The interrupt must be acknowledged BEFORE nesting starts. Reading ICC_IAR1 is what raises the GIC running priority to this interrupt's own, which masks it and everything of equal or lower priority; only then is re-enabling IRQ safe. My first attempt called nesting_start first and acknowledged inside the handler, so the still-pending, still-level-asserted timer was taken again the instant IRQ was enabled, and again, until the IRQ and System stacks were destroyed. It presented as garbage on the console and a hang with no fault to point at, and it broke demo_m3 and demo_threadx while leaving boot_check, demo_m2 and demo_mpu passing, because only the first two depend on the tick advancing. The Cortex-R5 example BSP states the requirement in one line: "ensure all IRQ interrupts are cleared prior to enabling nested IRQ interrupts." So entry.S now acknowledges in IRQ mode, carries the INTID in r4 -- which survives the mode switch, since only SP and LR are banked -- and also pushes it on the IRQ stack so a nested level reusing r4 cannot lose the outer level's value. End-of-interrupt waits until after nesting_end, in IRQ mode with interrupts masked, so dropping the running priority cannot re-admit the same interrupt. board_irq_handler splits in two. board_irq_service does the middle part on an already-acknowledged INTID and neither acknowledges nor EOIs; board_irq_handler keeps its old shape as the non-nesting entry point, so images built without TX_ENABLE_IRQ_NESTING behave exactly as before. The nesting instrumentation in irq_dispatch.c is inert unless an image asks for it. board_nest_provoke gates the SGI that the timer handler raises, so every other image sees the handler it always had. Verified on FVP_BaseR_AEMv8R. The nesting demo reports max depth 2, fifty nested SGIs across fifty ticks, depth unwound to zero, the tick still advancing, the low-priority thread still scheduled, and no spurious or unexpected INTIDs. In the same nesting configuration the five existing images all pass, so the split did not disturb the ordinary path. Both toolchains build it, GNU with no warnings and Arm Toolchain for Embedded 22.1.0 as well. Not done here: the S32Z280 example keeps its own irq_dispatch.c and entry.S and is untouched, so nesting on silicon is a separate change. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
b877305d13 |
Verified lazy VFP context switching on S32Z280 silicon (#609)
The FVP proved the lazy save and restore path at AR1/M5, but the board had never run it. Doing so needed one thing the model did not: the FPU has to be turned on. entry.S opens CPACR for CP10/CP11 and sets FPEXC.EN at EL1, guarded by __ARM_FP. EL2 already cleared HCPTR.TCP10/TCP11, but both of the EL1 gates read 0 out of reset, so a floating-point instruction raised an Undefined Instruction exception before this. tx_thread_vfp_enable() does not help: it sets the per-thread software flag that makes the context switch save and restore the registers, and never touches the hardware. Enabling it is the BSP's job. The __ARM_FP guard is what keeps a soft-float build assemblable, since "vmsr fpexc" is not a valid instruction for that target at all. demo_vfp_s32z280.c is the FVP's demo_m5.c test design over the LINFlexD console. The design is kept deliberately, because the two halves of the VFP context path need separate provocation: "fp check" holds eight live doubles across tx_thread_sleep, eight being enough to force the callee-saved D8-D15 bank that a solicited switch must preserve, while "fp busy" sits at the lowest priority and never sleeps, so the tick interrupts it mid-computation and exercises the interrupt half, D0-D15 plus FPSCR. Only these two threads opt in, so the test also shows that opting in is what does the work. s32z280_vfp.elf is gated on TX_R52_ENABLE_VFP, as demo_m5.elf is in the FVP example, since the image is meaningless unless the library was built with a floating-point ABI. Also made this example's -Wl,--no-warn-rwx-segments conditional on the compiler being GNU. #604 did that for the FVP example and this file still carried the literal, so ld.lld failed the link with "unknown argument". With that fixed the S32Z280 images build with Arm Toolchain for Embedded too. Verified on S32Z280 silicon: V2 D8-D15 bank preserved across switches PASS 50 solicited switches iterations = 0x000140A8 82,088 interrupted rounds corruptions = 0x00000000 V3 interrupted FP thread made progress PASS V4 no FP corruption across interrupts PASS filex_ptr = 0xF11EF11E V5 VFP flag did not alias filex_ptr PASS PASS lazy VFP context switch verified on silicon No regression: the soft-float build is warning-free and its image contains no vmsr at all, confirming the guard elides the block, and the existing demo still reports 100 ticks with 20 preemptions on the board. Both S32Z280 images also build with Arm Toolchain for Embedded 22.1.0, the hard-float VFP image included. No FVP file is touched. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
acdc02b2fc |
Assembled the code behind feature macros, and fixed the POP it found (#608)
scripts/check_clang.sh assembled every source with default flags, so the
preprocessor discarded each #ifdef block before the assembler saw it. Nothing in
the tree had ever assembled a guarded path. That covers the VFP context save and
restore in ten ports, and 218 files carrying TX_LOW_POWER or
TX_ENABLE_EXECUTION_CHANGE_NOTIFY.
Turning those on found a defect. The Cortex-M0 and Cortex-M23 execution-profile
paths bracket their call with
PUSH {r0, lr}
BL _tx_execution_isr_enter
POP {r0, lr}
and the last of those is invalid on Armv6-M and Armv8-M Baseline, where the
16-bit Thumb POP takes r0-r7 and pc and nothing else. GNU rejects it as well --
"cannot honor width suffix" -- so TX_ENABLE_EXECUTION_CHANGE_NOTIFY and
TX_EXECUTION_PROFILE_ENABLE have never been buildable on either port with either
toolchain. Four files, all the same shape.
The fix pops into a scratch register and moves it, MOV to a high register being
permitted where POP is not. r1 is free: the BL may clobber r0-r3, which is the
reason r0 is saved in the first place. Disassembling the result gives
push {r0, lr} / bl / pop {r0, r1} / mov lr, r1 / bx lr, one 16-bit instruction
more than before and otherwise the same.
Two findings that were not defects, recorded in the script so they are not
rediscovered:
Cortex-R4 needs an -mfpu to assemble its VFP path, because its FPU is an option
rather than part of the core. GNU fails identically without one, so this is a
flags requirement and not a toolchain divergence.
The A profile ports must not be given one. Adding -mfpu=vfpv3-d16 uniformly broke
28 files with "register expected", because those ports save D16-D31 and a -d16
FPU does not have those registers. Their defaults were already right.
The new stage runs under --asm-only as well, needing no target C library, and
reports 37 of 37 VFP files, 8 of 8 TX_LOW_POWER and 218 of 218
TX_ENABLE_EXECUTION_CHANGE_NOTIFY. Restoring the POP for one run makes it fail
with 217 of 218 and name the file and the error, so the stage is not vacuous.
The other four stages are unchanged: 711 of 711 assembled, 185 of 185 common C
sources for each of nine cores, 42 of 42 script-driven examples and 5 of 5 CMake
images.
The fixed code is verified to assemble with both toolchains and to encode as
intended. It is not verified running: there is no Cortex-M0 or Cortex-M23 model
here, and these are context save and restore paths, so that gap is worth stating.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
010a6c9fbb |
Gave the gnu ports a CMake build, which most of them lacked (#607)
The project guidelines ask for CMake and Ninja, but only 15 of the 59 gnu port directories had a CMakeLists.txt. None of the 27 AArch64 ports had one, so the architecture whose examples were repaired over the last few changes still could not be built the way the project says to build it, and nothing in CI could compile it. Add a CMakeLists.txt to the 44 that lacked one. Three of them are templates in ports_arch, because 34 of the 44 are generated: the ARMv7-A and AArch64 source lists are uniform within each family, so one template per family serves every core in it and update.sh distributes it. The other 10 ports have no generator and get their own file. Add the toolchain files those ports select, following the shape of cmake/cortex_a9.cmake. AArch64 needs a base file of its own rather than a variant of arm-none-eabi.cmake: it has no -marm or -mthumb to choose between and no -mfloat-abi, and aarch64-none-elf-gcc rejects -mlong-calls outright, so that flag cannot be carried across. The tools are named without a path, unlike cmake/cortex_r52.cmake which pins one, because pinning 30 files to a single machine's directory layout is the problem the previous change removed from the launch configurations. Three toolchain files cover ports that already had a CMakeLists.txt but no way to select it: the Armv8-M mainline gnu ports, cortex_m33, cortex_m55 and cortex_m85. Without cmake/<arch>.cmake the documented invocation cannot reach them. The top level needed one fix. It derives the SMP port directory as <arch>_smp, but ports_smp/linux and ports_smp/win64 predate that convention and carry no suffix, so those two could never be configured. Fall back to the bare name when the suffixed directory is absent. The check only fires when the suffixed directory does not exist, so no port that already resolved changes behaviour, and ports_smp/win64's existing CMakeLists.txt becomes reachable too. Verified by configuring and building every one: 53 of 53 static libraries build with cmake -G Ninja, using Arm GNU Toolchain 14.3.Rel1 for both arm-none-eabi and aarch64-none-elf. That covers the 44 new ports plus the 9 that already worked, and includes ports_smp/linux, which failed before the fallback. scripts/check_ports.sh passes, so the three templates and their 34 generated copies agree. Six gnu ports are still outside the CMake build, all for want of a compiler rather than a CMakeLists.txt: rxv1, rxv2 and rxv3 need the Renesas RX GNU toolchain and mips32_interaptiv_smp needs a MIPS one, neither of which is available here, so writing toolchain files for them would mean shipping untested guesses. risc-v32 and risc-v64 already build through their own differently named toolchain files. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
990d51b65d |
Brought up ThreadX on S32Z280 silicon and fixed the MPU encoding it exposed (#606)
* Added an S32Z280-594EVB example build for the Cortex-R52 port
First silicon bring-up for the Cortex-R52 port: an image that boots RTU0 core 0
on the NXP S32Z280-594EVB, drops from EL2 to EL1 and reports what the core says
about itself. Gated behind TX_R52_BUILD_S32Z280_EXAMPLE, separate from the FVP
example option so that FVP regression is unaffected by bring-up work and the
two boards' differing reset states cannot interact.
Verified on the board. The image runs from reset and reaches its completion
breakpoint, and the two CPSR values are the point of the exercise: mode 0x1A
(Hyp) at EL2 then 0x13 (Supervisor) at EL1, i.e. the EL2 to EL1 drop performed
by this code on silicon rather than on a model.
MIDR 0x411FD133 Cortex-R52 r1p3, part 0xD13
MPUIR 0x00001400 20 MPU regions at EL1
HMPUIR 0x00000014 EL2
CTR 0x8144C004
MPIDR 0x80000000
ID_PFR0 0x00000131
ID_PFR1 0x10111001 virtualisation field 1: EL2 implemented
SCTLR 0x70C50838 MPU, D-cache and I-cache all off at reset
CNTFRQ 0x00000000
The FVP reports MIDR 0x410FD0F0, part 0xD0F, which is an architecture envelope
model and not a Cortex-R52 -- so these are the first implementation-specific
numbers the port has had.
Three silicon facts the code exists to encode, each of which presents as
working hardware that quietly does the wrong thing:
- The core resets in THUMB state. CPSR reads 0x1FA out of reset because the
boot instruction NXP plants at the boot address is a T32 branch, so _start
is T32 and switches to A32 itself. An A32 entry would execute the first
halfword of its own instruction as Thumb.
- The debugger holds every core in debug state. NXP's
_reset_to_first_instruction() asserts MDM_AP CONTROL2[19:16] =
CR52_RTU0_{3,2,1,0}_EDBGREQ and never clears them, though its own comment
says start_debug_by_core_name() does. While asserted the core executes
nothing, yet registers and memory still respond and MC_ME/RGM report the
core released and clocked. tools/read_identity.gdb clears core 0's bit and
verifies the clear took.
- CNTFRQ reads zero, exactly as on the FVP, and is writable only at the
highest implemented exception level. entry.S records it rather than
writing a value, because the correct frequency for this board is not yet
established and a wrong one would silently mis-scale every derived
interval. Timer work must program it first.
Memory map: .text is linked and loaded at 0x79900000, the instruction-fetch
window and the reset address (MC_ME_PRTN0_CORE0_ADDR reads exactly that). It
is writable over the debug AXI port despite NXP's map declaring it read-only,
so no load-address alias is needed for a debugger-loaded image; a flash-booted
one would use the data alias at 0x32100000, which is the same physical SRAM as
confirmed by a sentinel write. Data, bss and the per-mode stacks go to
RTU-local data SRAM at 0x31780000. The TCMs are absent on purpose: they are
inaccessible at reset until their region registers are programmed, so nothing
needed for early boot can live there.
Reporting is built so that a failure cannot read as a success: boot_stage
distinguishes a fault from a hang and fault handlers record vector, syndrome
and address before parking; r52_identity.magic is written last so a
half-filled structure is detectable; and bsp_done() is a deliberate breakpoint
target so completion is observed rather than inferred from a timeout.
The image does not link ThreadX, so a failure here is unambiguously a boot or
board problem rather than a kernel one -- the same reasoning as the FVP's
boot_check.elf.
Not included: console, timer, GIC, MPU programming. LIN9 reaches the host
through the daughtercard USB-UART, which is LINFlex_9 at 0x42980000 and not
LINFlex_0; what is missing is the clock configuration needed to compute a baud
rate, and a console at the wrong rate produces garbage indistinguishable from a
crash. Hence reporting through memory for now.
FVP example builds and passes 6/6 unchanged.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Added a LINFlexD console to the S32Z280-594EVB example
Polled transmit console on LINFlex_9, which is the instance wired to the
daughtercard USB-UART through jumper J248. Not LINFlex_0: that is merely the
first instance in the Reference Manual's list and reaches no connector on this
board. bsp_boot.c now prints the identity registers as well as recording them
in memory; the memory copy stays, because it is what proved the boot path
before a console existed and it still works if the console is misconfigured.
The baud rate was derived rather than assumed, which mattered because a console
at the wrong rate produces garbage indistinguishable from a crash:
- LINFlex_9 is clocked by P5_LIN_BAUD_CLK, driven by MC_CGM_5 MUX 2.
- Read from the board: MUX_2_CSS selects source 2, MUX_2_DC_0 divides by 1.
- The EVB carries a 40 MHz crystal (UG10268 section 3.3.1.1), so 40 MHz.
- LFDIV = 40000000 / (16 * 115200) = 21.7014, giving LINIBRR 21 and
LINFBRR 11, for an actual 115274 baud -- 0.06% error.
Those are exactly the values the BootROM had already left in the registers,
which is the corroboration for the 40 MHz figure rather than merely a
plausible-looking calculation.
They are programmed here anyway instead of inherited, for two reasons.
Inheriting register state makes this image's behaviour depend on how the board
was last booted. And the BootROM's own configuration is wrong for a console:
it leaves PCE set and TxEn clear, because it is listening for a serial-boot
download rather than printing, so a host terminal on 8N1 would see framing
errors.
Verified on the board: the image reaches its completion breakpoint with the
console calls in the path, so every byte was accepted and UARTSR.DTF was set
for each -- the transmitter is configured, enabled and clocked. That does not
by itself prove the rate, since DTF sets at any baud; the rate rests on the
arithmetic above and the BootROM's independently matching divisors. Observing
the text needs the daughtercard USB-UART connected to a host terminal at
115200 8N1.
Register offsets and bit positions are from Reference Manual section 75.5.1
and the UARTCR/UARTSR diagrams. PCE at bit 2 and TxEn at bit 4 are called out
in the source because they are easy to transpose and the failure is silent.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Fixed two silent failures in the S32Z280 LINFlexD console
The console transmitted from the first attempt, but what arrived on the wire
was wrong in two independent ways, neither of which the target could detect.
Both are now fixed and the output is byte-for-byte correct, captured from the
USB-UART rather than read off a terminal.
1. UARTCR was written outside initialisation mode.
Setting LINCR1.INIT does not put the module in initialisation mode in the
same cycle, and most of UARTCR is writable only there. Configuring
immediately after the LINCR1 write half-worked: bits being *set* took
effect while bits being *cleared* did not. TxEn came on, but PCE stayed
on, so the line ran 8E1 against a host expecting 8N1. That corrupts only
those characters whose parity bit happens to be 0 and leaves the rest
readable, which looks like a marginal baud rate rather than a framing
error -- and would have sent the next hour into the clock tree.
linflexd_init() now polls LINSR[LINS] until the module reports
initialisation mode, reads UARTCR back afterwards, and returns a status
mask. bsp_boot.c records it and prints it as CONSOLE; 0 means both the
mode entry and every field write were confirmed rather than assumed.
2. DTF was cleared after the byte rather than waited on.
DTF is write-one-to-clear and does not de-assert in the same cycle as the
clearing write, so the next byte's poll could observe the previous byte's
flag, conclude the line was free while it was still busy, and have its own
write silently discarded. That cost exactly one character after every
"\r\n" pair -- the only place two bytes go out back to back -- so the
first letter of every line went missing: MIDR read as IDR, SCTLR as CTLR.
The transmit sequence is now write, wait for DTF, clear DTF, then wait for
the clear to take effect, with the same bounded guard as the init loop so
a stuck flag degrades to slow output rather than a hung boot.
Clearing before the write was tried and is worse, not better: the write
then lands while the previous byte is still shifting and is dropped, and
the poll afterwards sees the previous byte's completion, so most of the
output disappears. That failure is what established that the transmitter
discards writes while busy, which is the fact both bugs turn on. The
source says so, because the ordering looks arbitrary otherwise.
Verified end to end on the board, with the USB-UART passed through to WSL2 so
the received bytes could be compared against what the code intended to send:
=== ThreadX Cortex-R52 :: NXP S32Z280-594EVB ===
MIDR = 0x411FD133
MPUIR = 0x00001400
HMPUIR = 0x00000014
CTR = 0x8144C004
MPIDR = 0x80000000
ID_PFR0 = 0x00000131
ID_PFR1 = 0x10111001
SCTLR = 0x70C50838
CNTFRQ = 0x00000000
CPSR@EL2 = 0x000001DA
CPSR@EL1 = 0x600001D3
EL1 MPU regions: 00000014
CONSOLE = 0x00000000
=== boot complete ===
Reaching the completion breakpoint proved only that the transmitter ran; it
could not have caught either fault. Comparing bytes is what did.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Added generic timer support to the S32Z280-594EVB example
The Arm generic timer now runs on silicon: CNTFRQ is programmed at EL2 and a
one-shot compare fires when it should.
CNTFRQ programmed = 0x007A1200 (8000000)
elapsed counts = 8000043 (requested 8000000, +0.001%)
host interval = 1.0033 s
Both halves matter. The elapsed count shows the counter and the compare agree
with each other; the host interval shows the rate is actually 8 MHz rather than
merely self-consistent, which a wrong CNTFRQ would also produce. The residual
43 counts is about 5 us of polling overhead.
CNTFRQ reads zero out of reset here, exactly as on the Armv8-R AEM FVP -- so
that FVP behaviour is real silicon behaviour, not a model artefact, and the
port readme's warning to re-verify it was well placed. The parallel stops
there: the FVP also leaves the system counter itself stopped and its BSP must
start it, whereas on this board the BootROM has the counter running and only
the software-declared constant was missing. CNTFRQ is writable only at the
highest implemented exception level, so entry.S programs it before the drop to
EL1.
8 MHz was established three independent ways rather than assumed:
- Measured: CNTPCT sampled against host wall-clock time over a 32-second
interval gave 8.0227 MHz.
- Derived: RTU.GPR CFG_CNTDV reads 4, so the divider is (4+1) = 5, and the
board's FXOSC is 40 MHz -- itself already corroborated by the LINFlexD
baud divisors the BootROM left behind. 40 / 5 = 8.
- Confirmed: the one-shot compare above.
The compare is polled with the interrupt masked, deliberately. There is no
GIC configured yet, and proving the timer counts and fires first means a later
interrupt failure cannot be confused with the timer itself being wrong -- the
same staging that made the console tractable.
Also recorded, from a detour that cost a run: RTU0.GPR CFG_CNTDV at 0x76120010
is readable by the debugger, which returns 4, but a load from the core at EL1
kills the image. No fault handler runs, boot_stage stays at 4, and the debug
connection drops -- the signature of a stalled bus access rather than an abort,
and a failure mode that leaves nothing on-target to inspect. The debugger
reaches that window over the AXI-AP, which does not go through whatever gates
the core's own access to RTU peripheral space. The image no longer touches it;
the value lives in platform.h instead. This is the second address on this part
that is debugger-visible but not core-reachable, so peripheral windows are now
worth probing from the core before relying on them.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Enabled the R52 peripheral port and located the GIC on S32Z280
Two peripheral windows on this part stalled the core outright when read: no
abort, no fault handler, boot_stage frozen at its last value, and the debug
connection dropping with it. A stall gives you nothing on-target to inspect,
so both were found by printing a marker before each access and seeing which
marker came last. The Cortex-R52 TRM r1p3 (100026_0103_00_en) explains both;
NXP's reference manual mentions neither.
1. The low-latency peripheral port is disabled at reset.
IMP_PERIPHPREGIONR (TRM 3.3.80) describes a region at 0x76000000 of 4 MB
on this part -- the RTU peripheral space -- with separate enables for EL2
and EL1/0. Read from the board it was 0x76000034: base 0x76000000, size
0b01101 = 4 MB, and both enable bits clear. The TRM is explicit that each
"resets to 0". Until they are set, every access in that window stalls.
entry.S now sets both at EL2, which is also where it has to happen: EL1
writes to this register trap to EL2 when HACTLR.PERIPHPREGIONR is clear.
PERIPHPRG now reads 0x76000037 and the core reads RTU0.GPR CFG_CNTDV = 4
for itself -- the same value the debugger saw, which independently
confirms the divider behind the 8 MHz counter rate from the core's own
view rather than the debug path's.
2. The GIC is where NXP says, and needs an MPU mapping to reach.
IMP_CBAR (TRM 3.3.17) holds the physical base of the memory-mapped GIC
distributor in bits [31:21], its reset value wired from CFGPERIPHBASE.
Read from the board it is 0x47800000, confirming NXP's memory map from the
hardware rather than from a vendor debugger script -- the reference manual
contains no GIC base address anywhere. NXP's cryptic note that the "R52
Cluster set addr [31:21]" is just CFGPERIPHBASE[31:21] restated.
TRM Table 9-1 gives the frame layout relative to that base: distributor at
+0x000000, redistributor control at +0x100000, redistributor SGI/PPI at
+0x110000, then +0x20000 per further core. That matches what the FVP
example's driver already assumes, so it should port with address changes.
The distributor still stalls, and the TRM says why: "Ensure that the memory
region used for the GIC Distributor is configured as Device nGnRnE." The
MPU is disabled here -- SCTLR.M reads 0 -- so nothing maps that region.
Enabling the peripheral port does not help, the GIC being outside that
window and reached over AXIM. Reading it waits on MPU programming, which
is the next piece of work; the probe is removed rather than left in, since
while present it stalled every run and cost the console and timer results
that precede it.
gic_probe.c holds the two system-register reads. They cannot hang the bus,
unlike the memory accesses they diagnose, which is the point of them.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Reached the GIC on S32Z280 by mapping it Device nGnRnE
The GIC distributor and redistributor now respond on silicon:
MPU rgns = 0x00000003
SCTLR = 0x70C50839 (M set; caches still off)
GICD_PIDR2 = 0x0000003B architecture revision 3 = GICv3
GICD_TYPER = 0x0248001E
GICR_PIDR2 = 0x0000003B redistributor at base + 0x100000
Three independent sources had to agree to get here, and NXP's reference manual
supplied none of them. IMP_CBAR reports the distributor base from the hardware
(0x47800000). Cortex-R52 TRM Table 9-1 gives the frame layout: distributor at
+0x000000, redistributor control at +0x100000, SGI/PPI at +0x110000, then
+0x20000 per further core -- the layout the FVP example's gicv3.c already
assumes. And the TRM requires the region be Device nGnRnE, which was confirmed
rather than taken on faith: mapped Normal write-back the distributor stalls the
core outright, mapped Device it reads 0x3B.
mpu.c is ported from the FVP example, keeping its PRBAR/PRLAR encoding and the
reversed-AP-bit-order finding, with an S32Z280 region table. Caches are left
off: the point of this pass was reaching the GIC, and the debugger writes this
image straight into SRAM behind the caches, so enabling C and I deserves its own
step and its own check.
Diagnosis needed two new pieces of machinery, both kept:
- Memory-resident progress markers (probe_stage), mirroring the console
markers. Enabling the MPU can take the core down in a way that records no
fault AND takes the console with it, so boot_stage alone could not say how
far a risky sequence got. With markers inside mpu_init the failure went
from "somewhere in the MPU code" to "the SCTLR.M write itself" in one run.
- Resolving symbol addresses from the current ELF with nm on every
post-mortem. Adding probe_stage to the .data block shifted fault_vector,
so reusing addresses across builds reads the wrong words -- which nearly
produced a conclusion from a stale read.
OPEN QUESTION, recorded in mpu.c rather than papered over. A five-region map
that adds a second Device window for the RTU peripheral space (0x76000000, 4 MB)
makes the SCTLR.M write kill the core: all regions program without error,
probe_stage reaches 0x42, and the marker one instruction later never lands. No
exception is taken, so it stalls rather than aborts. Coverage is not the
explanation -- that map spanned the whole address space, and an earlier comment
of mine claiming otherwise was wrong and has been corrected. The cause is not
known. Consequence: RTU peripheral access after the MPU is enabled is untested;
the CFG_CNTDV read this image does happens before the enable.
Also not tightened yet: code is writable and data executable in this map. Wrong
for a protection demo, right for bring-up where an unmapped address costs a
silent stall rather than a fault. Narrowing it belongs with the ThreadX
integration, which can verify it.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Fixed an out-of-bounds MPU table write that made enabling the MPU stall the core
The five-region map now programs correctly and the MPU enables with the GIC
reachable:
R idx PRBAR PRLAR
R 0 00000000 477FFFC1
R 1 47800002 479FFFC3 GIC, XN, Device nGnRnE
R 2 47A00000 75FFFFC1
R 3 76000002 763FFFC3 RTU peripherals, XN, Device nGnRnE
R 4 76400000 FFFFFFC1
MPU rgns = 5, SCTLR = 0x70C50839 (M set)
GICD_PIDR2 = 0x3B, GICD_TYPER = 0x0248001E, GICR_PIDR2 = 0x3B
The bug was mine and it was simple: mpu_regions[] is declared with a fixed
size, and the FVP example this file came from sized it [3] to match the three
regions it programs. A five-region table wrote mpu_regions[3] and [4] past the
end of the array.
What made it hard to see is what the corruption produced. Every region
appeared to program without error, and the failure was a stall on the SCTLR.M
write -- no abort, no fault handler, no exception, and the debug connection
dropping. Reading the regions back out of the hardware, with the MPU
deliberately left disabled so the readback could not itself stall, showed
region 3 with base 0x00000000 instead of 0x76000000. With its correct limit
that region spanned 0x00000000-0x763FFFFF and overlapped regions 0 to 2, and
PMSAv8-R leaves overlapping regions UNPREDICTABLE -- so stalling was permitted
behaviour rather than a hardware fault.
One bug accounts for every observation: three regions worked, five failed, and
it failed identically whether the extra window was Device or Normal, because
the memory type was never involved.
Fixes: the table capacity is now named (MPU_TABLE_REGIONS) and sized 16, and
mpu_init checks mpu_regions_used against it as well as against MPUIR. The
existing check only compared against the hardware's region count, which said 20
and told us nothing about the array.
Two earlier claims of mine, corrected here rather than left standing. A comment
asserting "narrow coverage was the problem" was wrong: the failing map covered
the whole address space. And this is NOT a defect inherited from the FVP
example -- that code declares [3] and uses three regions, entirely
self-consistent. Worth offering upstream only as hardening: the bare [3] with
no bound check is what let a growing map overrun it silently.
The region readback is kept in the image. It costs eight lines a boot and
verifies the map against the hardware every time, which is precisely what
turned an unexplained stall into a one-line diagnosis.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Enabled interrupt-driven ticks by clearing SCTLR.TE
Timer interrupts are delivered, acknowledged and re-armed on silicon:
SCTLR = 0x30C50838 -> 0x30C50839 (TE cleared, M set)
IRQ count = 7
timer INTID = 30 measured, not assumed
spurious = 0
unexpected = 0
SCTLR.TE -- Thumb Exception enable -- RESETS SET on this part. Every vector
table in entry.S is A32, so the core entered each exception in T32 state,
decoded an A32 branch as Thumb, landed mid-instruction at a misaligned address
and took an undefined-instruction exception, which repeated the same way. No
handler ever ran. entry.S now clears it at EL1 and clears HSCTLR.TE at EL2.
The consequence was worse than a crash: it made every fault invisible.
boot_stage stayed at its last value, fault_vector stayed zero, and the core sat
at el1_vectors+4. A data abort from an unmapped peripheral, an abort from an
overlapping MPU region, and a correctly delivered timer interrupt all presented
identically -- an unexplained hang with nothing recorded. What gave it away was
reading the banked registers: LR_irq and SPSR_irq showed an IRQ had been taken
from SVC with interrupts enabled, CPSR showed UND mode with T set, and LR_und
pointed at a misaligned address inside the A32 vector table.
The Armv8-R AEM FVP resets with TE clear, which is why the same vector tables
work there and why nothing in the FVP work anticipated this.
This commit also corrects earlier comments of mine. Several places described
these failures as "a stalled bus access rather than an abort", with the absence
of a recorded fault as the evidence. That was wrong: they were ordinary
exceptions whose handlers were unreachable. The two underlying bugs -- the
peripheral port disabled at reset, and the out-of-bounds MPU table write -- were
real and are correctly fixed, but the explanation for why they presented so
mutely was not. Fixed in platform.h, mpu.c and bsp_boot.c.
Interrupt plumbing: gicv3.c is ported from the FVP example with the frame bases
in platform.h, derived from IMP_CBAR and TRM Table 9-1. el1_irq_entry in
entry.S uses the classic A32 form rather than ThreadX's context save/restore,
since this image does not link the kernel. irq_dispatch.c counts ticks instead
of calling _tx_timer_interrupt, and keeps separate spurious and unexpected-INTID
counters so that "no interrupt arrived" and "an interrupt arrived and was
mishandled" cannot be confused. timer_start_oneshot_irq arms with IMASK clear;
the polled timer_start_oneshot keeps IMASK set, and the two are separate
functions because either mistake is silent.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Ran ThreadX on S32Z280 silicon with threads, tick and preemption
The kernel runs on the board:
=== ThreadX Cortex-R52 :: S32Z280-594EVB ===
console = 0x00000000
entering kernel
=== ThreadX on S32Z280: results ===
ticks = 0x00000064 100 ticks
sleeper = 0x00000014 20 wakeups
spinner = 0x00159B48 1415496 iterations
preempt = 0x00000014 20 preemptions
PASS threads, tick and preemption all verified
The figures are internally consistent, which is what makes them evidence rather
than encouragement: 100 ticks elapsed for a 100-tick sleep, so the tick rate is
exactly as configured; 20 sleeper wakeups is precisely 100/5 for a 5-tick sleep;
and 20 preemptions means every wakeup took the core from a runnable
lower-priority thread rather than receiving it by cooperative handoff.
This validates the context-switch assembly merged in #579 on real Cortex-R52
silicon. Until now it had only run against the Armv8-R AEM FVP, which reports
MIDR part 0xD0F -- an architecture envelope model, not an R52 implementation.
The demo judges four things separately because on first silicon they fail
independently: that time advanced (the GIC delivers the timer PPI and
_tx_timer_interrupt is reached), that tx_thread_sleep returned (tick-driven
scheduling, not merely the interrupt), that the lowest-priority thread ran (the
sleeper actually yielded), and that the sleeper resumed while the spinner was
runnable (preemption). A demo that only counted whether both threads ran would
pass with a broken tick if they happened to yield to each other.
Integration pieces, all from the FVP example with the board's differences:
- tx_initialize_low_level.S publishes the system stack and the first free
address, then calls board_init(). Shorter than the A-profile reference
ports because entry.S has already given every mode its own stack from
dedicated linker regions.
- entry.S routes the IRQ vector through _tx_thread_context_save and
_tx_thread_context_restore under TX_R52_USE_THREADX_IRQ, keeping the
standalone A32 handler for s32z280_boot.elf, which does not link the kernel
so that a boot failure there stays unambiguous.
- board_init() runs mpu_init() FIRST. The GIC distributor is unreachable
until its region is mapped Device nGnRnE, so any GIC access before the MPU
is enabled aborts.
- link.lds provides _end as well as end; _tx_initialize_low_level publishes
unused memory from _end.
- bsp_done() is defined in the demo rather than shared from bsp_boot.c, whose
bsp_main would collide with the demo's.
s32z280_boot.elf still builds and the FVP example still passes 6/6.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Tightened the S32Z280 MPU map and verified protection is enforced
The map now carries real permissions -- code read-only and executable, data
writable and never executable, peripherals Device nGnRnE and never executable,
everything else unmapped -- and the enforcement is demonstrated rather than
asserted:
X1 write to read-only code region
faults = 0x00000001
fault DFSR = 0x00000A0C bit 11 (WnR) set: a write fault
fault addr = 0x79900000 the address written
X2 PASS write to code faulted and was recovered
Exactly one fault, identified as a write, at the precise address, followed by
successful resumption. ThreadX still passes on the same map: ticks 100,
sleeper 20, spinner ~1e6, preempt 20, and IRQ count keeps advancing during the
protection test.
Reading SCTLR back or listing the programmed regions would only show what was
configured. Provoking the violation shows it is enforced, which is the claim
that matters for a protection story.
entry.S gains a recoverable data-abort path, gated on a fault_expected flag the
test arms. It records DFSR and DFAR, counts the fault, and resumes at the
instruction after the faulting access -- lr on data-abort entry is the faulting
address plus 8, so subs pc, lr, #4 skips the access instead of retrying it
forever. Unarmed, a data abort stays fatal and reported, which is what a real
bug should get.
This supersedes the permissive bring-up map, and the reason that map existed is
worth recording: a narrow map had failed earlier and the cause was unknown, so
coverage was blamed. Coverage was never the problem. An out-of-bounds write to
a [3]-sized region table put a region at base 0 overlapping everything, and
SCTLR.TE being set meant the resulting abort could not reach a handler. With
both fixed a narrow map works first time, and -- more usefully -- a mistake in it
now reports vector, syndrome and faulting address instead of hanging mutely.
Until TE was cleared no fault handler on this board could run, so no protection
claim about it could be tested at all. This is the first one that could.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Enabled the S32Z280 caches, with an honest effectiveness result
Both caches are on and verified from SCTLR:
CLIDR = 0x09200003
SCTLR = 0x30C50839 -> 0x30C5183D (C and I set)
cachesOn = 1
counts off = 0x00094DF4
counts on = 0x00094CC8
gain/1000 = 0
C3 PASS caches enabled (SCTLR.C and SCTLR.I set)
C4 no significant speedup -- expected here, the workload is already
in fast local SRAM
cache.c supplies what the FVP example explicitly deferred to silicon: a
data-cache set/way sweep that reads CLIDR and CCSIDR and walks every set and
way the hardware reports, rather than assuming a geometry. The FVP file notes
that this "belongs with the silicon bring-up where it can be verified against
the real cache geometry"; this is that. The sweep matters because caches are
not architecturally guaranteed invalid out of reset, and enabling a write-back
data cache holding stale valid lines would evict them over live memory.
The effectiveness result is reported separately from the enable, and that
separation was earned. The first version of this test passed on warm < cold and
printed "caches on and the workload got faster" for 609904 counts against
609686 -- a 0.036% difference that is measurement noise. That criterion was
worthless and the PASS was misleading. It now reports the gain as a fraction
and only claims a speedup above 10%.
No speedup here is a plausible result rather than a defect: both the code and
the data this workload touches live in RTU-local low-latency SRAM, close to core
speed already. There is no slow memory on this board to demonstrate a cache
against -- DDR would be the place, and DDR is not initialised. So the honest
claim is "enabled and harmless", not "enabled and beneficial".
cache_clean_all() is called before parking, and this is not optional. With a
write-back data cache, values this image writes can sit dirty in cache where a
debugger reading SRAM cannot see them -- and the memory-resident progress
markers and fault records this bring-up depends on are read exactly that way.
Without the clean, a post-mortem read could report stale values and look like a
fault that never happened.
ThreadX still passes on the same configuration and the protection test still
faults correctly: write to the read-only code region takes exactly one write
fault at the written address and recovers.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Made the cache benchmark actually measure the cache
The caches now show a real effect instead of noise:
D line = 64 bytes
D ways = 4
D sets = 64
D bytes = 0x4000 16 KB L1 data cache, read from CCSIDR
bench wds = 0x800 8 KB working set, half the cache
counts off = 0x000D1B87 858503
counts on = 0x0009F515 652053
gain/1000 = 0xF0 24% faster
C4 PASS caches measurably faster (>=10%)
The previous benchmark could not have measured anything, and its 0.036%
"speedup" was noise. Two things were wrong with it.
It used a static buffer in the RTU-local fast-data bank. That memory is close
to core speed already, so caching it saves almost nothing -- there is no latency
to hide. The benchmark now runs over the extended SRAM at S32Z_EXT_SRAM_BASE,
which the Reference Manual describes as NOT RTU-local, and mpu.c gains a sixth
region to map it Normal write-back and never executable.
And its working set was a guessed 4 KB constant. It is now sized at run time
from CCSIDR to half the reported data cache, so it fits and is revisited every
pass. A set larger than the cache would stream through and evict, showing
little benefit even over slow memory -- a guess could have produced a null
result for a reason that has nothing to do with whether the cache works.
24% on this loop is a believable figure rather than a suspiciously large one:
the workload is a store-heavy read-modify-write over a write-back cache, so part
of the cost is write traffic that still has to reach SRAM.
Useful by-product: the L1 data cache geometry is now read and printed rather
than assumed. Nothing in the SoC reference manual gives it, and the set/way
invalidate sweep in cache.c depends on it being right.
ThreadX still passes and the protection test still faults correctly.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Fixed the PRBAR field encoding and proved execute-never enforcement
program_region() shifted every PRBAR field one bit too far left: SH<<4,
AP<<2 and XN<<1, where the Cortex-R52 TRM (r1p3 figure 3-39, table 3-80)
places BASE[31:6], RES0[5], SH[4:3], AP[2:1] and XN[0]. The intended XN
therefore landed in AP[1] -- the EL0-access bit -- and no region was ever
execute-never. On this board the core ran instructions straight out of
.data with LR_abt at 0x3178000c to prove it.
The write-to-read-only test passed throughout, which is why this survived:
under the old shift the AP value's low bit happened to land in the real
AP[2], the read-only bit, so permissions came out right by accident while
XN was silently discarded. Only an instruction fetch from a region marked
non-executable could distinguish the two.
The AP macros in mpu.h go back to the architectural encoding, AP[2] for
read-only and AP[1] for EL0 access. No calibration is needed; the values
were correct as published all along.
Verified on S32Z280 silicon: the data region now reads back PRBAR
0x31780001 with XN set, a write to the read-only code region faults with
DFSR 0xA0C, execution from the data region takes a prefetch abort with
IFSR 0x20C at 0x31780000, both faults recover, and the ThreadX demo still
reports 100 ticks with 20 preemptions.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Corrected the same PRBAR encoding bug in the FVP MPU example
The FVP example carries the identical off-by-one shift that the S32Z280
bring-up exposed, so no region there was execute-never either: the data
region and the whole peripheral region were both freely executable, and
each also gained unintended EL0 access from the misplaced XN bit.
This also retires the "PRBAR.AP bit order is reversed" claim in mpu.h.
That note rested on a real measurement -- four regions, one per AP
encoding, privileged write attempted on each, writes faulting for 0b01 and
0b11 -- but the cause was the shift, not the bit order. With AP written
into bits[3:2], its low bit lands in the real AP[2], the read-only bit, so
writes fault exactly when that bit is set. The "high bit grants EL0
access" half of the conclusion was never tested; under the old shift that
bit landed in SH[0], programming a shareability the TRM calls
UNPREDICTABLE. The architectural encoding needs no calibration.
demo_mpu.c gains the check that would have caught this: an instruction
fetch from the execute-never data region must take a prefetch abort.
Recovering from one cannot work the way the data-abort path does, since
skipping the faulting instruction is impossible when the instruction is
what could not be fetched, so entry.S gains a recoverable prefetch-abort
path that returns to a landing point recorded by mpu_try_execute.
Verified both ways. All five FVP images pass, and the new check reports
prefetch aborts 0 -> 1. Restoring the old shift for one run makes it fail
with 0 -> 0, so the check is not vacuous.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Removed the unassemblable .arch directive from the S32Z280 example
GNU as accepts ".arch armv8-r"; LLVM's integrated assembler accepts no spelling
of it -- not armv8-r, not armv8r, not armv8-r+crc -- and stops with "Unknown
Arch: armv8-r". There is nothing to substitute, so the directive goes. The
architecture comes from -mcpu=cortex-r52 on the command line, which every
toolchain file passes, so it only restated it.
entry.S in this example never had one. The two equivalents in the FVP example are
handled separately, in the change that brings those images under the LLVM check.
Verified: both S32Z280 assembly sources now assemble with Arm Toolchain for
Embedded 22.1.0, and the GNU build of s32z280_boot.elf and s32z280_demo.elf is
unchanged with no warnings.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
* Stopped read_identity.gdb reporting failure after a successful demo run
The script reads r52_identity, which only the boot image defines. The kernel demo
shares the same bsp_done breakpoint and the same boot_stage marker, so it is
convenient to point the same script at either image -- but against the demo the
lookup raised and gdb exited 1, after the demo had already printed
"PASS threads, tick and preemption all verified" over the console. A tool that
reports failure on success is worse than one that prints nothing.
The lookup is now guarded, and says which image it is looking at rather than
falling over. Everything after it is skipped when the structure is absent.
Verified against both images on S32Z280 silicon: the demo exits 0 and reports its
own PASS, and the boot image still prints the full identity report with MIDR
0x411FD133 and C3, C4, X2 and X4 all passing.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
b0ad67bae9 |
Brought the Cortex-R52 examples under the LLVM check, and fixed two blockers (#604)
#600 added a line reporting which example builds the LLVM check passes over, and cortex_r52 was on it: its examples are driven by CMake rather than by a build_threadx.sh pair, so the linking stage never touched them. Covering them turned up two reasons they could not have been built with anything but GNU. .arch armv8-r has no portable spelling. GNU as accepts it, and LLVM's integrated assembler rejects every variant -- armv8-r, armv8r, armv8-r+crc -- with "Unknown Arch: armv8-r". There is nothing to substitute, so the directive is gone from entry.S and tx_initialize_low_level.S; -mcpu=cortex-r52 already selects the architecture, both toolchain files pass it, and the directive only restated it. Worth noting where this hid: the assembly stage walks ports/*/gnu/src, so example assembly had never been assembled by LLVM at all. -Wl,--no-warn-rwx-segments is GNU ld only, added in binutils 2.39. ld.lld does not warn about RWX segments and rejects the flag outright, failing the link with "unknown argument". It is now selected on CMAKE_C_COMPILER_ID rather than spelled into all six targets, and the reason it exists at all -- a bare-metal image has one flat DRAM region and leaves access control to the MPU -- moves to the one place that sets it. cmake/cortex_r52_clang.cmake is the toolchain file. It names the tools as found on PATH, which is what CI uses, then pins $HOME/toolchains if that directory exists, mirroring how cortex_r52.cmake pins the GNU toolchain and for the same reason. Falling back rather than requiring the pinned path keeps the file usable on a machine that keeps clang elsewhere. THREADX_TOOLCHAIN stays "gnu": there is no clang port directory, this builds the gnu sources with a different compiler, which is what the whole check does for every other Arm port. check_clang.sh gains a fourth stage for CMake-driven examples, and no longer reports cortex_r52 as a gap. It reads the image list out of the generated ninja graph rather than repeating it, so adding a target cannot escape the check, and filters out the cmake_object_order_depends_target_* phonies -- counting those reported ten images where there are five. Verified with Arm Toolchain for Embedded 22.1.0, the version the workflow pins. All five images link, and all five then run and pass on FVP_BaseR_AEMv8R: boot_check, demo_m2, demo_m3, demo_threadx and demo_mpu. That is a step beyond the AArch64 examples, which are link-verified only. The full check reports 711 of 711 assembly sources, 185 of 185 common C sources for each of nine cores, 42 of 42 script-driven examples and 5 of 5 CMake images, leaving only cortex_a5_smp, cortex_a7_smp and cortex_a9_smp listed as having no driver. GNU is unaffected: the same five images build with no warnings and the FVP test suite passes 5 of 5. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
31fbe21d20 |
Removed developer machine paths from the shipped project files (#603)
The Arm Development Studio launch configurations and the IAR project files
carried absolute paths from the machines they were last opened on: four
individuals' home directories, a Broadcom support build tree, and directories
such as C:\release\threadx, C:\temp1702\tx and D:\threadx. They are published in
the repository and none of them resolves for anyone else.
In the launch files all of it is saved session state, which files that were never
opened in a debugger simply do not have:
breakpoints 72 files, saved breakpoints anchored to source
paths under cortex_a35 on one machine. Emptied
rather than deleted, matching the 27 launch files
that already carry an empty list.
scripts_view_script_links 70 files, a scripts view cache. The script the
launch actually runs is named one line earlier
through ${workspace_loc:...}, which is portable.
substitutePath 65 files, source lookup remapping. 30 map a path
to itself, and the rest point at one machine's
copy of libgloss or a build directory.
TREE_NODE_PROPERTIES 39 files, which rows the variables view had
expanded.
DebugCommandLine.History 2 files, the debug console command history,
including a typed absolute path.
The two IAR cases are not session state and are corrected rather than removed:
IarchiveOutput 23 files. The librarian output path pointed at
another machine, so the library was written
outside the project. Set to ###Unitialized###,
which 86 other option blocks already use, letting
IAR derive $EXE_DIR$\$PROJ_FNAME$.a.
IlinkIcfFile 1 file. The linker configuration file is load
bearing, so the file name is kept and the path
made project relative, as 33 other option blocks
spell it.
Seven of the files are ports_arch templates, so the generators would otherwise
have copied the paths back.
Verified: all 162 launch and IAR project files in the tree parse as well-formed
XML afterwards, no tracked launch or project file mentions any of the four user
names, and scripts/check_ports.sh passes, so the templates and their generated
copies still agree.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
c78c080139 |
Made the ARMv8-A debug launch configurations name their own core (#602)
Two identifiers in the Arm Development Studio launch configurations still said A35 in every AArch64 port, because the generator's patch table could not reach them. The config database taxonomy id is spelt in lower case in the launch files, /platform/armfvp/base_a35x4, but the search string read base_A35x4. Both generators are case sensitive here, sed in update.sh and -creplace in update.ps1, so the rule matched nothing and all 27 ports kept the A35 taxonomy id while the platform name and model beside it named the right core. That the rule exists at all shows the substitution was intended, and the ARMv7-A generator does the same thing correctly, which is why its ports carry a per-core ve_cortex_a<n>x1 id. The SMP launch files name the activity "Cortex-A35x4 SMP" where the ThreadX ones say "Debug Cortex-A35", so the existing activity rule reached the ThreadX ports only and every SMP port advertised an A35 activity. Add a rule that matches the " SMP" suffix, which also keeps it from rewriting the FVP model name that the line above already handles. Both changes are mirrored in update.ps1, which stays equivalent to update.sh. Verified by regenerating: 50 launch files change and nothing else, 100 lines of taxonomy id across 24 ThreadX and 26 SMP files, and 52 lines of activity name across the 26 SMP files. No other attribute in those files moves, so the new " SMP" rule does not touch the model name. The two Cortex-A35 ports are untouched, as they must be, since substituting A35 for A35 is a no-op. scripts/check_ports.sh passes, so the result is reproducible. This cannot be exercised here: confirming that Arm Development Studio resolves the per-core taxonomy ids needs Arm Development Studio. The change rests on the platform name and model already naming the core, on the patch rule's existence, and on the ARMv7-A ports having shipped per-core ids all along. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
96b1638fe4 |
Gave the Armv8-M examples a linkable sample and fixed what stopped them (#601)
The Cortex-M23, M33, M55 and M85 examples each shipped a build_threadx.sh with no build_threadx_sample.sh, so scripts/check_clang.sh skipped all four at its linking stage: the whole Armv8-M family had no example-link coverage. Adding the missing script surfaced five separate reasons why none of them could have linked. ports/cortex_m23/gnu/src/tx_initialize_low_level.S referenced Image$$ARM_LIB_STACK$$ZI$$Limit and __Vectors, which are Arm toolchain scatter-load names that no GNU linker defines. This is a GNU port, so it now uses __RAM_segment_used_end__ and _vectors, as the Cortex-M33 version already does. The Cortex-A ports mention Image$$ZI$$Limit only inside a comment; this was a live relocation. The Cortex-M23 crt0 needed unified assembler syntax. The Armv7-M file it derives from carries no .syntax directive, so GNU as reads it in the legacy divided syntax, where a Thumb data-processing instruction sets the flags whether or not the mnemonic says so, and "mov r2, #0" assembles to 2200, MOVS. LLVM implements unified syntax only, where those mnemonics mean the non-flag-setting forms that Armv8-M Baseline does not have. Disassembling the Cortex-M4 object confirms GNU already emits 2200 movs, 1a52 subs and 3001 adds, so spelling them out changes no encoding. The flags carry meaning: crt0_memory_copy branches on the result of "subs r2, r2, r1" and again on "subs r2, #1". Cortex-M23 has no SVC_Handler to name in its vector table. Its library is built -DTX_SINGLE_MODE_NON_SECURE, and tx_thread_schedule.S defines that handler only when neither TX_SINGLE_MODE_SECURE nor TX_SINGLE_MODE_NON_SECURE is set, so the SVCall slot takes __tx_BadHandler. The table also uses the CMSIS handler names throughout, because that is what the Armv8-M ports export: the Armv7-M table references __tx_SVCallHandler and __tx_SysTickHandler, which resolve to nothing here. The other three needed a C library. tx_thread_secure_stack.c calls malloc and free, which pulls the allocator in, and the sample scripts already defined SYSCALL_LIB for exactly that without ever passing it to the linker. Their linker scripts now also provide end and _end, which newlib's libnosys _sbrk wants, and __heap_start and __heap_end, which picolibc wants instead. Cortex-M55 and M85 build with the hard-float ABI now, in both the library and the sample. Arm Toolchain for Embedded ships no soft-float MVE multilib and says so plainly: "No library available for MVE with soft-float ABI." The library and the sample have to agree, and -mfloat-abi=hard is what check_clang.sh already uses for these two cores in PORT_TARGET. Verified with GNU 13.2.1 and Arm Toolchain for Embedded 22.1.0. Every port links under both: Cortex-M23 at 206,136 and 312,636 bytes, M33 at 305,856 and 320,696, M55 at 306,228 and 320,664, M85 at 306,240 and 320,668. check_clang.sh now reports 42 of 42 example builds linking, up from 38, and no longer lists any port as carrying build_threadx.sh without build_threadx_sample.sh. The images are link-verified only and have not been executed, as with the AArch64 examples. Only .sh drivers are added. Cortex-M23 has a build_threadx.bat with no sample counterpart and the other three have no .bat at all; adding untested Windows scripts belongs in its own change. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
b87a62d218 |
Brought Cortex-A34 and the A78 SMP port under the ARMv8-A generator (#599)
ports/cortex_a34, ports_smp/cortex_a34_smp and ports_smp/cortex_a78_smp were absent from the ARMv8-A generator's core list, so ports_arch never reached them and scripts/check_ports.sh could not detect that they had drifted. They had. The two Cortex-A34 ports sat two releases behind the shared sources, at 6.1.10 against 6.3.0. The difference is not only banners: tx_initialize_low_level.S lacked the SUB x1, x1, #15 that precedes the BIC when the system stack pointer is recorded, so the value stored was the incoming SP rather than the first 16-byte boundary below it. The Arm Development Studio example was missing GICv3_aliases.h, which every other port's GICv3_gicc.h includes to reach the interrupt controller registers through the stringify indirection. The Cortex-A78 SMP port had never had its debug launch configuration patched at all. It still named the A35 platform and model, so Debug Cortex-A35, Base_A35x4 and FVP_Base_Cortex-A35x4 would have started an A35 model for an A78 port. Add cortex_a34 to the core list. Cortex-A78 cannot go in the same list, because the generator walks cores against port sets and would then create a ports/cortex_a78 that the tree has never had; give it a separate SMP-only list consulted only for the tx_smp port set. Both changes are mirrored in update.ps1, which stays equivalent to update.sh. Verified by regenerating: exactly these three port directories change, 116 files modified and 13 added, and no already-covered core moves, so the core list was the only thing holding them back. scripts/check_ports.sh passes, which now means these three are checked for reproducibility for the first time. scripts/check_clang.sh reports 38 of 38 example builds linked, the three new ports included. readme_threadx.txt is not added here. Only the two Cortex-A35 template ports carry it; the other 23 AArch64 ports do not, so its absence from Cortex-A34 is the tree's norm rather than drift, and supplying it everywhere is separate work. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
51ca412b36 |
Added build scripts for the AArch64 examples, which had none (#597)
The AArch64 examples could only be built through the Arm Development Studio project files beside them. There was no script, so nothing in CI or on a developer's machine could build them, and the sources they need were never named anywhere: vectors.S, v8_aarch64.S and v8_utils.S were present but unreferenced, which is why a hand written link failed on GetCPUID, GetAffinity, InvalidateUDCaches and ZeroBlock. Those four are defined in v8_aarch64.S and v8_utils.S, in the tree all along. Add one pair of scripts, in ports_arch so every AArch64 port receives them. Both derive the -mcpu value and the kernel source directory from the port directory they sit in, so a single pair serves the twelve ThreadX ports and the twelve SMP ports, the latter building against common_smp and picking up the C sources those ports carry alongside the assembly. TOOLCHAIN=atfe selects Arm Toolchain for Embedded in place of the GNU toolchain, as it does for the ARMv7 example scripts. One symbol genuinely had no definition. startup.S calls initialise_monitor_handles to open the standard file handles over a debugger connection; the GNU toolchain provides it in libgloss through --specs=rdimon.specs, while picolibc has no equivalent and neither does the LLVM toolchain's semihosting library. semihost_stub.S supplies a weak no-op, linked for that toolchain only, so a real definition always wins. Extend scripts/check_clang.sh to cover ports_smp as well as ports, since the example builds now exist there too. Verified by building every AArch64 example with Arm Toolchain for Embedded: twelve ThreadX ports at 314,744 to 315,256 bytes of text and twelve SMP ports at 325,304 to 325,880. scripts/check_clang.sh now reports 35 of 35 example builds linking with none failing, the twenty-four new ones plus the eleven that already built. The images have not been executed: the sample targets the Base platform peripheral addresses, which the available emulator does not provide. Cortex-A34, and the Cortex-A34 and A78 SMP ports, are left out. They are not in the generator's core list, so they receive nothing from ports_arch and cannot be checked for reproducibility either. Bringing them in rewrites 114 files in those three directories, which deserves its own change. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
08eff8061c |
Made the Cortex-A12, A15 and A17 examples link with an LLVM toolchain (#596)
Those three sample scripts compiled crt0.S and reset.S but named neither on the link line, and omitted -nostartfiles, so the toolchain was also free to link its own startup. Under GNU that produced a working image; under an LLVM toolchain picolibc's crt0 was pulled in and wanted __data_start, __data_source, __data_size, __bss_size, __stack, __tls_base and __arm32_tls_tcb_offset, none of which the linker script defines. Defining picolibc's contract in the shared linker script was tried first and abandoned: it reaches into ABI constants that cannot be verified here, and it is unnecessary, because the in-tree crt0.S is already a complete startup for this script. It sets up the stack and zeroes BSS between __bss_start__ and __bss_end__, and .data carries no AT() so there is nothing to copy. So the three scripts now pass -nostartfiles and name crt0.o and reset.o, which is what the four sibling A profile cores already do and what these scripts were evidently compiling those files for. Both the old and the new GNU images contain the in-tree startup, __vectors from reset.S and _mainCRTStartup from crt0.S, so the startup was already being used through implicit startup file resolution rather than an explicit operand. How that resolution happened without the objects being named is not accounted for here, which is itself the argument for naming them: the same implicit behaviour does not hold across toolchains, and that is why the LLVM link failed. Add a readme to each of the three example directories. Their tx_initialize_low_level.S is the generic ARMv7-A skeleton, 300 lines and identical across all three, with no interrupt controller programming, no timer and no vector table installation, where the A5, A7, A8 and A9 examples have all three and ship the matching Versatile Express support files. These three therefore demonstrate that the port builds; they will not receive a timer tick. Nothing in the tree said so, which invites the assumption that they are equivalent. Verified with both toolchains for all three cores. GNU links at 41,276 bytes of text, down from 41,768 because the unused toolchain startup is no longer included, and Arm Toolchain for Embedded links at 37,982 where it previously could not link at all. The resulting image is structurally sound: entry at _start, __vectors at address zero, and a BSS range in RAM. It has not been executed. scripts/check_clang.sh now links eleven of eleven example builds. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
a8aebc5206 |
Completed the Cortex-M0 barriers and made its example build with both toolchains (#595)
Two unrelated Cortex-M0 gaps, both left over from earlier work. The memory barriers from #523 reached the Cortex-M0 gnu port but not its ac6 or iar siblings, in either the inline system return in tx_port.h or the assembly routine. Both take the same GNU or IAR code path, and iar already carried the entry barrier, so the missing pieces were the entry pair for ac6 and the barrier after restoring the interrupt posture for both. All three tools now match. The Cortex-M0 example could not link with any toolchain. cortexm0_crt0.S references 23 linker script symbols and the script defined only 12 of them, so __text_start__, __text_end__, __text_load_start__, the rodata and fast section symbols, and the ctors and dtors load addresses were all unresolved. The Cortex-M4 script defines all 23, including a .fast section with no content whose symbols exist so that the startup copy is a no-op, and its comment says as much. The Cortex-M0 script is brought to that same shape. That left the example failing under LLVM only, on instructions that ARMv6-M can encode just one way. The file declared .code 16 but no syntax mode, so GNU as used the legacy divided syntax in which a plain add or sub sets the flags implicitly, while LLVM implements unified syntax only and rejected the non-flag-setting spelling. Declaring .syntax unified and writing movs, adds and subs makes both assemblers agree, and the encodings GNU produces are byte identical before and after, verified by disassembling both objects. The Cortex-M0 example now links with GNU at 22,520 bytes of text and with Arm Toolchain for Embedded at 22,866, so it comes off the list of examples not expected to link and scripts/check_clang.sh now links eight of eight. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
f3ab36dacc |
Made the example builds work with LLVM, and fixed four that were broken on Linux (#594)
* Made the example builds work with LLVM, and fixed four that were broken on Linux The gnu example builds are the natural LLVM path: Arm Toolchain for Embedded is LLVM based, consumes GNU ld linker scripts, and needs none of the scatter files or Arm DS projects the ac6 examples carry. So rather than port the ac6 examples, teach the gnu scripts to drive either toolchain. Each script now selects the toolchain from a TOOLCHAIN variable that defaults to gnu, so every existing invocation behaves exactly as before. TOOLCHAIN=atfe switches the compiler, adds the target triple, passes the entry symbol through to the linker rather than to the driver, and links the toolchain's own semihosting library in place of --specs=nosys.specs, which is a GCC spec file mechanism with no LLVM equivalent. That last choice avoids adding a syscall stub source to every example. The scripts are parameterised rather than duplicated. Copies of build scripts would drift the first time one side was edited, which is the failure this repository has just spent several changes recovering from. Four sample scripts could not run on a case-sensitive filesystem at all. They compiled MP_PrivateTimer.s and V7.s while the files on disk are MP_PrivateTimer.S and v7.s, and the Cortex-A5 and A9 scripts named MP_GIC.s where the file is MP_GIC.S. The link lines named V7.o accordingly. Corrected to match the files, which is why the Cortex-A5, A7, A8 and A9 examples now build on Linux where before they could not. Extend scripts/check_clang.sh to link the example builds as well as compile the sources, since compiling proves the sources parse while only linking exercises entry symbols, linker scripts and the C library together. Eight examples are listed as not expected to link, each with its reason, so the gaps stay visible rather than being silently skipped. Five of them fail with the GNU toolchain too and are therefore not LLVM problems: the Cortex-M0 example's crt0 references __text_load_start__, __text_start__ and __text_end__, which its linker script never defines, while the Cortex-M4 script defines the equivalents; the arm9, arm11, Cortex-R4 and Cortex-R5 examples need newlib multilib variants that are not present in every GNU toolchain packaging. The Cortex-A12, A15 and A17 examples fail only with LLVM, because their link line omits -nostartfiles so the toolchain's own crt0 is linked and wants picolibc's __data_start, __data_source, __data_size and __bss_size, which their linker script does not define. Verified by building every example with both toolchains. Seven link with both: Cortex-A5, A7, A8, A9, M3, M4 and M7. The GNU results are unchanged where they worked before, and now also succeed for the four scripts with the case bug. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> * Resolved the compiler path before running the example builds The example stage runs each build script from inside its own directory, so a relative path passed with --clang stopped resolving there and every example build failed instantly. Local runs passed an absolute path and did not show it; the CI job passes a path relative to the workspace root, which did. Resolve the compiler to an absolute path once, before any directory change, and print it so the toolchain in use is visible in the log. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> * Parameterised the archiver as well, and stopped the check hiding failures The toolchain selection covered the compiler but not arm-none-eabi-ar, which the scripts invoke 628 times to assemble the library archive. A machine with an LLVM toolchain and no GNU one therefore could not build any example, which is exactly the situation in CI: the runner has no Arm GNU toolchain installed. Local runs had one, so this only appeared once the job ran. Select the archiver alongside the compiler, taking llvm-ar from beside clang in the toolchain. The four scripts that call arm-none-eabi-ld directly are left alone: they are arm9, arm11, Cortex-R4 and Cortex-R5, all already listed as not expected to link, and their invocations are specific to GNU ld in ways that parameterising would not resolve. The check reported "example build produced no image" and then filtered the log for lines containing "error", which hid the actual cause, since a missing tool reports "command not found" or "No such file or directory". It now prints the tail of the log. That filtering cost two CI round trips to diagnose something the first run already knew. Verified by shadowing arm-none-eabi-gcc, arm-none-eabi-ar, arm-none-eabi-ld and the aarch64 equivalents with stubs that fail loudly, then running the whole check: all 711 sources assemble, all eight profiles compile and all seven example builds link without any GNU tool being invoked. The GNU default path still produces an identical image. The diagnostics were confirmed by pointing the archiver at a name that does not exist and checking that the reason appears. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
68eb205768 |
Made the Arm ports build with LLVM, and added a check that keeps them that way (#593)
The gnu ports are only ever built with GNU tooling, and GNU as accepts several non-canonical forms that LLVM's assembler rejects. Nothing noticed, because nothing built them with anything else. This matters beyond clang itself: Arm Toolchain for Embedded is LLVM based and is the successor to Arm Compiler 6, so these are the code paths ac6 users move onto. Seven files needed changing, none of which alters the emitted code: LDREX and STREX take no offset in A32 state; the #imm form is Thumb-2 only. GNU as drops the redundant zero, LLVM rejects it. Removed from the Cortex-A5, A7 and A9 SMP protect routines. ARMv8-M Baseline has no flag-preserving MOV immediate, so GNU as already emits MOVS. Writing MOVS in the two Cortex-M23 sources says what the assembler was doing anyway. One of them sits in a branch only compiled for the single mode secure configurations, which is why it had never surfaced. The Cortex-M0 schedule routine wrote LDR r0, =#0x10000000 with a stray hash, which its own sibling file already wrote correctly. The Cortex-M0 system return routine selected the numbered subsection .text 32, which makes LLVM place the constant pool beyond the range a Thumb-1 PC relative load can reach. Plain .text fixes it and GNU accepts either form. The reason is recorded in the file, since 32 files pair a numbered subsection with a literal pool load and the rest only escape because Thumb-2 and A32 have far more range. Add scripts/check_clang.sh, which assembles every Arm gnu port source and compiles the common C sources for one core per architecture profile, and a clang_check workflow that installs Arm Toolchain for Embedded and runs it. The toolchain version is pinned and checksum verified, for the same reason the runner image is pinned. The port directory to target mapping in that script is explicit rather than prefix matched. Prefix matching is what makes cortex_a5 also match cortex_a53 and cortex_a55, which are AArch64, and assembling those as ARM32 produces hundreds of misleading errors; that mistake cost real time while measuring this, so the reason is recorded next to the table. Verified with Arm Toolchain for Embedded 22.1.0: 711 of 711 assembly sources assemble and 185 of 185 C sources compile for all eight profiles, against 6 assembly failures before the change. arm-none-eabi-gcc still assembles all 315 ARM32 sources, so nothing regressed for GNU. The check was confirmed to fail when any one of the fixes is reverted. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
f23a1807a2 |
Added bash A profile update scripts and restored ports_arch as their source (#592)
The ARMv7-A and ARMv8-A ports are generated by update.ps1, which needs PowerShell, so the cortex-a job in ports_arch_check ran on a Windows image and nobody could reproduce it locally on Linux. Add update.sh beside each update.ps1, with the same cores, compilers, copy sets and patches, and move the job to the same Linux image as everything else. The bash scripts were checked against the PowerShell ones by comparing what each reports as drifted. They agree exactly on the 63 files the Windows job last reported, and differ on 12 more, which turn out to be a defect in update.ps1 rather than in the port. Its two .cproject patterns are written as 'value=`"cortex-a7`"' with backticks that survive into the pattern, so that replacement has never matched, while the neighbouring Cortex-A7.NoFPU pattern has no backticks and always worked. The result is that the AC6 example builds for the A5, A8, A9, A12, A15 and A17 cores name cortex-a7 as their CPU while their FPU string is correct. The bash scripts do what the PowerShell ones intended, so regenerating corrects those twelve files. Restore ports_arch as the source for the rest. The implementation of _tx_thread_smp_time_get from #555 was applied to the twenty four generated SMP ports and never to ports_arch, which still held MOV x0, #0 with a FIXME comment, so regenerating would have replaced a working generic timer read with a stub. That implementation now lives in the source. The remaining differences are cosmetic and resolve in favour of the source: a trailing blank line in 38 copies of tx_thread_schedule.S and comment spacing in one tx_port.h. Note that the Cortex-A VFP fix is already present in ports_arch and was never at risk, contrary to what the description of the port consistency checks change said before this was measured. Extend scripts/check_ports.sh to run the A profile generators too, and make it fail when a generator fails or is missing rather than reporting a clean tree, which would have been a false pass. Pin every workflow to ubuntu-24.04. ubuntu-latest already resolves to that image, so nothing changes today, but a future migration becomes a deliberate commit rather than something that happens underneath the -m32 builds. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
eb4ec4e3b5 |
Restored ports_arch as the source of truth for the Cortex-M ports (#590)
The Cortex-M ports under ports/ are generated. scripts/copy_armv7_m.sh copies one tx_port.h and the per tool sources to fifteen M3, M4 and M7 targets, and scripts/copy_armv8_m.sh does the same for nine M33, M55 and M85 targets. The ports_arch_check workflow runs both scripts and fails if the tree is not reproducible, so those copies are meant never to be edited directly. They were. Every Cortex-M fix since #523 was applied to the generated copies and not to the source, so the source fell behind and the check went red: running the three scripts on dev changes 35 files. The check triggers only on pull requests targeting master, which is why nothing caught it while the fixes were merged into dev. Left alone, the next run of these scripts would have reverted three separate pieces of work: the memory barriers and clobbers from #523, the correction of the IAR assembly header to use the assembler's own comment syntax, and the move of tx_initialize_low_level.S into example_build for the M33, M55 and M85 GNU ports from #514. Bring the sources up to what the ports carry today, and regenerate. Two behavioural changes come with that, both deliberate. The barriers from #523 reach the ac5 and keil variants of M3, M4 and M7, which were outside the scope of that fix and never received it. The barrier that follows restoring the interrupt posture, which #523 gave only to the GNU ports because GNU was the only toolchain that could be tested, now applies to every tool; the identical asm statement already shipped in the AC6 and IAR ports, so this adds a pipeline flush rather than any new compiler exposure. Regenerating also drops a stray #endif at the end of the Cortex-M85 IAR tx_port.h, added by #523, which left that header with one more #endif than #if and unable to compile. Every other ARMv8-M port was balanced. Verified that the scripts are idempotent afterwards, that ports_arch_check would pass, that no port loses a barrier or a clobber, that every regenerated header is preprocessor balanced, and that every Cortex-M port covered by the two scripts now carries the entry barrier. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
cbdf924d6b |
Removed the duplicated function body in the Cortex-M4 AC6 port (#589)
_tx_thread_system_return_inline() in the Cortex-M4 AC6 tx_port.h was followed
by a second, orphaned copy of its own body. The copy had no function header, so
it declared interrupt_save at file scope and then placed statements there,
which does not compile. It is also the older version of the body, without the
dsb and isb barriers, so it was left behind rather than intended: the barriers
were added by commit
|
||
|
|
f84dd3d2aa |
Added Cortex-R52 port with Armv8-R AEM FVP example build (#579)
Added a ThreadX port for Arm Cortex-R52 with the GNU toolchain
Adds ports/cortex_r52/gnu together with a CMake build, an example build for the
freely available Armv8-R AEM FVP, and an automated test suite.
The port is EL2-aware by construction. Cortex-R52 always implements EL2 and
resets into it, so the reset path configures EL2, installs both the EL2 and EL1
vector tables, and only then drops to EL1 to run the kernel. Keeping that
structure from the start lets partitioning work reuse the boot path rather than
replace it; TX_R52_BOOT_AT_EL1 skips the EL2 stage where a vendor monitor has
already dropped privilege.
This is also the first R-profile port in the tree with a working CMake build.
cmake/cortex_a9.cmake exists but ports/cortex_a9/gnu/CMakeLists.txt is empty, so
no A- or R-profile port could be built this way before.
Contents:
- Kernel port: 16 assembly sources plus tx_port.h, seeded from the
Cortex-R5/GNU port.
- Build: cmake/cortex_r52.cmake and the port CMakeLists.txt, with soft/hard
float, VFP, FIQ and IRQ/FIQ nesting options.
- Example BSP: EL2 to EL1 boot, GICv3, generic timer, PMSAv8-R MPU,
semihosting and PL011 consoles.
- Tests: six images registered with CTest, each judging itself and
terminating the model.
- tx_port_offset_check.c: compile-time assertions on the TX_THREAD offsets
the assembly reaches by hard-coded displacement. This is the same defect
class fixed for Cortex-R4/R5 in #578; nothing in the toolchain ties those
literals to the C structure.
Two facts were established by experiment rather than taken from documentation,
and are recorded in the code and the readme because both are traps:
- PRBAR.AP bit order is reversed relative to th
AArch64 macro set: the low bit is read-only and the high bit grants EL0
access. Programming four disjoint regions, one per encoding, and
attempting a privileged write to each gave 0b
and 0b10 allowed. Using the AArch64 ordering
as read-only and silently accept writes, because region coverage is still
enforced: an unmapped address faults while a "read-only" region does not.
- The generic timer PPI is INTID 30, recorded from ICC_IAR1 after enabling
the whole PPI range 25-31, rather than assume
Model behaviour that a silicon port must revisit: the FVP leaves CNTFRQ at zero
with the system counter stopped, so the BSP programs both; CNTHCTL.PL1PCTEN,
CNTHCTL.PL1PCEN and ICC_HSRE must be set at EL2 or EL1 accesses trap; and
tx_thread_vfp_enable() sets only a per-thread software flag, so enabling the FPU
hardware (CPACR, FPEXC.EN) is the board support package's responsibility, not
the kernel's. The FVP reports MIDR 0x410FD0F0, an architecture envelope model
rather than a Cortex-R52, so it validates architecture and not implementation.
Nothing in the port is gated on MIDR.
Changes outside ports/cortex_r52, three and all small: cmake/cortex_r52.cmake is
a new toolchain file; CMakeLists.txt gains enable_testing() at the top level,
without which add_test() in a subdirectory generat
the root never references, so ctest reports no tests from the build root; and
.gitignore covers build directories and __pycache__.
Verification: all six images build warning-free a
FVP_BaseR_AEMv8R in three configurations -- protection off, protection and
caches on, and both together with hard-float VFP. Six build-option
combinations build clean and pass the tick-and-preemption demo at run time.
The tests assert behaviour rather than configuration, which is what caught the
problems above: the cooperative demo asserts the order of execution, because
counting alone cannot distinguish working context switches from one thread
running to completion, and the MPU test provokes real faults, because reading
SCTLR back only proves a bit was set.
The 96-test regression suite is host-side and pas
configurations on the Linux port. It is not cross-run here: it validates
portable kernel logic, while the port-specific risk is the assembly, which the
FVP images exercise. Structural coverage is therefore not claimed.
Deliberate deviations: passing a string literal to tx_thread_create reports a
discarded const qualifier from common/inc/tx_api.
CHAR *; demo_threadx.c keeps the shipped sample's int main() so the demo body
stays byte-identical to it; and the floating-point test compares exactly on
purpose, since a tolerance would mask a restored
wrong.
Not included: modules support, split-mode SMP and
support. The example targets the FVP only, and NXP S32Z280 bring-up is
separate work; the readme flags what to re-verify there.
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
|
||
|
|
2baae7c57f |
Fixed VFP issues on Cortex-R4 and R5 ports (#578)
* Fixed missing VFP thread extension on Cortex-R4 and R5 ports Building cortex_r4/gnu, cortex_r5/gnu or cortex_r5/ac6 with TX_ENABLE_VFP_SUPPORT corrupts memory. The assembly in tx_thread_schedule, tx_thread_system_return and tx_thread_context_restore reads and writes a per-thread VFP enable flag as [thread, #144], but every TX_THREAD_EXTENSION in these ports' tx_port.h was empty, so no such member existed. Offset 144 in the resulting TX_THREAD is tx_thread_filex_ptr, so the damage runs both ways: - tx_thread_vfp_enable() stores 1 into tx_thread_filex_ptr, after which FileX dereferences 0x1; - conversely, once FileX sets that pointer the context switch reads it as "floating point enabled" and starts pushing about 132 bytes of floating-point context onto a thread stack never sized for it, then restores unrelated data into the D registers. sizeof(TX_THREAD) is 180 on these ports, so 144 is a live member rather than padding past the end of the structure. What makes this dangerous is that it is silent. The defect was reproduced at runtime by reverting the Cortex-R52 port -- which inherited the same assembly and the same hard-coded 144 from the R5 port -- to this state and running its floating-point demo on the Armv8-R AEM FVP. Every floating-point value was still preserved exactly and no corruption was reported across interrupts, because a non-zero filex_ptr reads as "enabled" and the context switch dutifully saves and restores. The only visible damage was tx_thread_filex_ptr reading 0x00000001. In other words the feature you enabled appears to work, and what breaks is an unrelated pointer that nothing notices until FileX is introduced. With the field present the same demo reports the sentinel intact. This appears to be drift rather than a deliberate limitation. All fourteen Armv7-A port and toolchain combinations define the field and contain the VFP code paths. The R-profile family is inconsistent in both directions: these three have the code paths without the field, while cortex_r4/ac5, cortex_r4/ac6, cortex_r4/iar, cortex_r5/ghs, cortex_r5/iar, cortex_r7/ghs and cortex_r8_smp/ac5 define the field but contain no VFP code paths. The fix follows the Armv7-A ports exactly: TX_THREAD_EXTENSION_2 carries tx_thread_vfp_enable, and tx_thread_vfp_enable/disable are declared. The field is defined unconditionally rather than under TX_ENABLE_VFP_SUPPORT because the offset is hard-coded in assembly, so a conditional member would shift every following field and be correct in only one configuration. No assembly changes are needed: 144 is already correct once the member exists. Verified with GCC 14.3: on all three ports tx_thread_vfp_enable now lands at exactly offset 144, matching the assembly, with tx_thread_filex_ptr moved to 148. The library builds cleanly for cortex_r4/gnu and cortex_r5/gnu both with and without TX_ENABLE_VFP_SUPPORT, and the VFP save and restore instructions appear in the port assembly only when it is requested. cortex_r5/ac6 is verified by offset and inspection only, as that toolchain was not available here. Runtime verification was performed on Armv8-R as described above rather than on R4/R5 silicon; upstream QEMU's xlnx-zcu102 machine holds its Cortex-R5F cores in reset, so no R5 target was available. ABI note for release documentation: adding the member moves tx_thread_filex_ptr and everything after it by four bytes and grows TX_THREAD from 180 to 184 bytes. This is the layout the Armv7-A ports already have, so the change aligns R-profile with A-profile rather than introducing a new one, but kernel awareness and TraceX tooling that hard-codes offsets will need rebuilding. Follow-up worth considering separately: nothing in the toolchain ties the literal 144 in the assembly to the C structure, which is what allowed this to go unnoticed. A compile-time assertion per port turns any recurrence into a build failure -- when the experiment above reverted the field, that assertion failed the build before a corrupting binary could be produced. ports/cortex_r52/gnu/src/tx_port_offset_check.c is a working reference. * Added compile-time offset checks to the Cortex-R4 and R5 ports Nothing in the toolchain connects the structure offsets hard-coded in the port assembly to the C definition of TX_THREAD, which is what allowed the missing VFP thread extension to go unnoticed. These files assert the offsets that the context-switch path depends on, so a recurrence becomes a build failure instead of memory corruption discovered later at runtime. Three offsets are asserted: the VFP enable flag at 144, the thread stack pointer at 8 and the run counter at 4. The stack pointer and run counter precede every extension macro in TX_THREAD, so they are stable by construction. Offset 144 was checked against the build options that could plausibly move it -- stack checking, event trace, event logging, performance info and TX_NOT_INTERRUPTABLE -- and is unchanged by all of them, because the only conditional member nearby guards tx_thread_filex_ptr, which follows the extension, and TX_THREAD_USER_EXTENSION appears much later in the structure. The assertions therefore cannot misfire on a legitimate configuration. C99 has no _Static_assert, so a negative array dimension is used. Each file compiles clean under -std=c99 -pedantic -Wall -Wextra and emits zero bytes of code or data. Verified that the checks actually catch regressions rather than merely compiling: asserting a deliberately wrong offset is rejected on all three ports, and removing the VFP field again fails the build outright with a message naming the missing member. These ports have no CMake build of their own, so the files take effect only in builds that compile everything under src/. That is still worthwhile given they cost nothing and emit nothing. Extending the same technique to the remaining ports is deliberately left as separate work: it requires reading each port's assembly to attribute every hard-coded literal to the right structure, and copying assertions without that analysis would risk asserting wrong offsets. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com> |
||
|
|
a33f6efc48 |
Updated Cortex-R4 Thumb bit immediates (#565)
Use explicit immediate syntax when testing the Thumb bit in AC6 Cortex-R4 stack build assembly. Co-authored-by: Codex <codex@openai.com> |
||
|
|
3fcfb4e4c9 |
Updated Cortex-M BASEPRI zero immediates (#564)
Use explicit immediate syntax when clearing BASEPRI in GNU and AC6 Cortex-M scheduler assembly. Co-authored-by: Codex <codex@openai.com> |
||
|
|
d9e7dfea89 |
Add SysTick counter reset in Cortex-M ports (#561)
* Add SysTick counter reset in Cortex-M ports The SysTick counter value (SYST_CVR @0xE000E018) is indeterminate after reset. Without clearing it prior to enabling the counter, the first tick interval becomes unpredictable. * Fixed missing comment and added generated-by header in Cortex-M0/AC6 port Added the missing inline comment on the LDR r1, =0 instruction (Build value for SysTick reset) and the required AI-generated attribution comment under the copyright header. --------- Co-authored-by: Frédéric Desbiens <frederic.desbiens@eclipse-foundation.org> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> |
||
|
|
859c098747 |
Fix nested comment warning in RX GCC ports (#550)
The RX GCC assembly port contains an unclosed C-style comment in tx_thread_context_save.S. Because .S files are passed through the C preprocessor before assembly, the nested comment can trigger GCC warnings. Close the affected comment block without changing assembly instructions or runtime behavior. Signed-off-by: FranCDoc <fchiesadoc@gmail.com> |