Commit Graph
131 Commits
Author SHA1 Message Date
Frédéric Desbiens 010a6c9fbb Gave the gnu ports a CMake build, which most of them lacked (#607)
The project guidelines ask for CMake and Ninja, but only 15 of the 59 gnu port
directories had a CMakeLists.txt. None of the 27 AArch64 ports had one, so the
architecture whose examples were repaired over the last few changes still could
not be built the way the project says to build it, and nothing in CI could
compile it.

Add a CMakeLists.txt to the 44 that lacked one. Three of them are templates in
ports_arch, because 34 of the 44 are generated: the ARMv7-A and AArch64 source
lists are uniform within each family, so one template per family serves every
core in it and update.sh distributes it. The other 10 ports have no generator
and get their own file.

Add the toolchain files those ports select, following the shape of
cmake/cortex_a9.cmake. AArch64 needs a base file of its own rather than a
variant of arm-none-eabi.cmake: it has no -marm or -mthumb to choose between and
no -mfloat-abi, and aarch64-none-elf-gcc rejects -mlong-calls outright, so that
flag cannot be carried across. The tools are named without a path, unlike
cmake/cortex_r52.cmake which pins one, because pinning 30 files to a single
machine's directory layout is the problem the previous change removed from the
launch configurations.

Three toolchain files cover ports that already had a CMakeLists.txt but no way
to select it: the Armv8-M mainline gnu ports, cortex_m33, cortex_m55 and
cortex_m85. Without cmake/<arch>.cmake the documented invocation cannot reach
them.

The top level needed one fix. It derives the SMP port directory as
<arch>_smp, but ports_smp/linux and ports_smp/win64 predate that convention and
carry no suffix, so those two could never be configured. Fall back to the bare
name when the suffixed directory is absent. The check only fires when the
suffixed directory does not exist, so no port that already resolved changes
behaviour, and ports_smp/win64's existing CMakeLists.txt becomes reachable too.

Verified by configuring and building every one: 53 of 53 static libraries build
with cmake -G Ninja, using Arm GNU Toolchain 14.3.Rel1 for both arm-none-eabi
and aarch64-none-elf. That covers the 44 new ports plus the 9 that already
worked, and includes ports_smp/linux, which failed before the fallback.
scripts/check_ports.sh passes, so the three templates and their 34 generated
copies agree.

Six gnu ports are still outside the CMake build, all for want of a compiler
rather than a CMakeLists.txt: rxv1, rxv2 and rxv3 need the Renesas RX GNU
toolchain and mips32_interaptiv_smp needs a MIPS one, neither of which is
available here, so writing toolchain files for them would mean shipping
untested guesses. risc-v32 and risc-v64 already build through their own
differently named toolchain files.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-12 12:43:57 -04:00
Frédéric Desbiens 990d51b65d Brought up ThreadX on S32Z280 silicon and fixed the MPU encoding it exposed (#606)
* Added an S32Z280-594EVB example build for the Cortex-R52 port

First silicon bring-up for the Cortex-R52 port: an image that boots RTU0 core 0
on the NXP S32Z280-594EVB, drops from EL2 to EL1 and reports what the core says
about itself.  Gated behind TX_R52_BUILD_S32Z280_EXAMPLE, separate from the FVP
example option so that FVP regression is unaffected by bring-up work and the
two boards' differing reset states cannot interact.

Verified on the board.  The image runs from reset and reaches its completion
breakpoint, and the two CPSR values are the point of the exercise: mode 0x1A
(Hyp) at EL2 then 0x13 (Supervisor) at EL1, i.e. the EL2 to EL1 drop performed
by this code on silicon rather than on a model.

    MIDR      0x411FD133   Cortex-R52 r1p3, part 0xD13
    MPUIR     0x00001400   20 MPU regions at EL1
    HMPUIR    0x00000014   EL2
    CTR       0x8144C004
    MPIDR     0x80000000
    ID_PFR0   0x00000131
    ID_PFR1   0x10111001   virtualisation field 1: EL2 implemented
    SCTLR     0x70C50838   MPU, D-cache and I-cache all off at reset
    CNTFRQ    0x00000000

The FVP reports MIDR 0x410FD0F0, part 0xD0F, which is an architecture envelope
model and not a Cortex-R52 -- so these are the first implementation-specific
numbers the port has had.

Three silicon facts the code exists to encode, each of which presents as
working hardware that quietly does the wrong thing:

  - The core resets in THUMB state.  CPSR reads 0x1FA out of reset because the
    boot instruction NXP plants at the boot address is a T32 branch, so _start
    is T32 and switches to A32 itself.  An A32 entry would execute the first
    halfword of its own instruction as Thumb.

  - The debugger holds every core in debug state.  NXP's
    _reset_to_first_instruction() asserts MDM_AP CONTROL2[19:16] =
    CR52_RTU0_{3,2,1,0}_EDBGREQ and never clears them, though its own comment
    says start_debug_by_core_name() does.  While asserted the core executes
    nothing, yet registers and memory still respond and MC_ME/RGM report the
    core released and clocked.  tools/read_identity.gdb clears core 0's bit and
    verifies the clear took.

  - CNTFRQ reads zero, exactly as on the FVP, and is writable only at the
    highest implemented exception level.  entry.S records it rather than
    writing a value, because the correct frequency for this board is not yet
    established and a wrong one would silently mis-scale every derived
    interval.  Timer work must program it first.

Memory map: .text is linked and loaded at 0x79900000, the instruction-fetch
window and the reset address (MC_ME_PRTN0_CORE0_ADDR reads exactly that).  It
is writable over the debug AXI port despite NXP's map declaring it read-only,
so no load-address alias is needed for a debugger-loaded image; a flash-booted
one would use the data alias at 0x32100000, which is the same physical SRAM as
confirmed by a sentinel write.  Data, bss and the per-mode stacks go to
RTU-local data SRAM at 0x31780000.  The TCMs are absent on purpose: they are
inaccessible at reset until their region registers are programmed, so nothing
needed for early boot can live there.

Reporting is built so that a failure cannot read as a success: boot_stage
distinguishes a fault from a hang and fault handlers record vector, syndrome
and address before parking; r52_identity.magic is written last so a
half-filled structure is detectable; and bsp_done() is a deliberate breakpoint
target so completion is observed rather than inferred from a timeout.

The image does not link ThreadX, so a failure here is unambiguously a boot or
board problem rather than a kernel one -- the same reasoning as the FVP's
boot_check.elf.

Not included: console, timer, GIC, MPU programming.  LIN9 reaches the host
through the daughtercard USB-UART, which is LINFlex_9 at 0x42980000 and not
LINFlex_0; what is missing is the clock configuration needed to compute a baud
rate, and a console at the wrong rate produces garbage indistinguishable from a
crash.  Hence reporting through memory for now.

FVP example builds and passes 6/6 unchanged.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Added a LINFlexD console to the S32Z280-594EVB example

Polled transmit console on LINFlex_9, which is the instance wired to the
daughtercard USB-UART through jumper J248.  Not LINFlex_0: that is merely the
first instance in the Reference Manual's list and reaches no connector on this
board.  bsp_boot.c now prints the identity registers as well as recording them
in memory; the memory copy stays, because it is what proved the boot path
before a console existed and it still works if the console is misconfigured.

The baud rate was derived rather than assumed, which mattered because a console
at the wrong rate produces garbage indistinguishable from a crash:

  - LINFlex_9 is clocked by P5_LIN_BAUD_CLK, driven by MC_CGM_5 MUX 2.
  - Read from the board: MUX_2_CSS selects source 2, MUX_2_DC_0 divides by 1.
  - The EVB carries a 40 MHz crystal (UG10268 section 3.3.1.1), so 40 MHz.
  - LFDIV = 40000000 / (16 * 115200) = 21.7014, giving LINIBRR 21 and
    LINFBRR 11, for an actual 115274 baud -- 0.06% error.

Those are exactly the values the BootROM had already left in the registers,
which is the corroboration for the 40 MHz figure rather than merely a
plausible-looking calculation.

They are programmed here anyway instead of inherited, for two reasons.
Inheriting register state makes this image's behaviour depend on how the board
was last booted.  And the BootROM's own configuration is wrong for a console:
it leaves PCE set and TxEn clear, because it is listening for a serial-boot
download rather than printing, so a host terminal on 8N1 would see framing
errors.

Verified on the board: the image reaches its completion breakpoint with the
console calls in the path, so every byte was accepted and UARTSR.DTF was set
for each -- the transmitter is configured, enabled and clocked.  That does not
by itself prove the rate, since DTF sets at any baud; the rate rests on the
arithmetic above and the BootROM's independently matching divisors.  Observing
the text needs the daughtercard USB-UART connected to a host terminal at
115200 8N1.

Register offsets and bit positions are from Reference Manual section 75.5.1
and the UARTCR/UARTSR diagrams.  PCE at bit 2 and TxEn at bit 4 are called out
in the source because they are easy to transpose and the failure is silent.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Fixed two silent failures in the S32Z280 LINFlexD console

The console transmitted from the first attempt, but what arrived on the wire
was wrong in two independent ways, neither of which the target could detect.
Both are now fixed and the output is byte-for-byte correct, captured from the
USB-UART rather than read off a terminal.

1.  UARTCR was written outside initialisation mode.

    Setting LINCR1.INIT does not put the module in initialisation mode in the
    same cycle, and most of UARTCR is writable only there.  Configuring
    immediately after the LINCR1 write half-worked: bits being *set* took
    effect while bits being *cleared* did not.  TxEn came on, but PCE stayed
    on, so the line ran 8E1 against a host expecting 8N1.  That corrupts only
    those characters whose parity bit happens to be 0 and leaves the rest
    readable, which looks like a marginal baud rate rather than a framing
    error -- and would have sent the next hour into the clock tree.

    linflexd_init() now polls LINSR[LINS] until the module reports
    initialisation mode, reads UARTCR back afterwards, and returns a status
    mask.  bsp_boot.c records it and prints it as CONSOLE; 0 means both the
    mode entry and every field write were confirmed rather than assumed.

2.  DTF was cleared after the byte rather than waited on.

    DTF is write-one-to-clear and does not de-assert in the same cycle as the
    clearing write, so the next byte's poll could observe the previous byte's
    flag, conclude the line was free while it was still busy, and have its own
    write silently discarded.  That cost exactly one character after every
    "\r\n" pair -- the only place two bytes go out back to back -- so the
    first letter of every line went missing: MIDR read as IDR, SCTLR as CTLR.

    The transmit sequence is now write, wait for DTF, clear DTF, then wait for
    the clear to take effect, with the same bounded guard as the init loop so
    a stuck flag degrades to slow output rather than a hung boot.

    Clearing before the write was tried and is worse, not better: the write
    then lands while the previous byte is still shifting and is dropped, and
    the poll afterwards sees the previous byte's completion, so most of the
    output disappears.  That failure is what established that the transmitter
    discards writes while busy, which is the fact both bugs turn on.  The
    source says so, because the ordering looks arbitrary otherwise.

Verified end to end on the board, with the USB-UART passed through to WSL2 so
the received bytes could be compared against what the code intended to send:

    === ThreadX Cortex-R52 :: NXP S32Z280-594EVB ===
    MIDR      = 0x411FD133
    MPUIR     = 0x00001400
    HMPUIR    = 0x00000014
    CTR       = 0x8144C004
    MPIDR     = 0x80000000
    ID_PFR0   = 0x00000131
    ID_PFR1   = 0x10111001
    SCTLR     = 0x70C50838
    CNTFRQ    = 0x00000000
    CPSR@EL2  = 0x000001DA
    CPSR@EL1  = 0x600001D3
    EL1 MPU regions: 00000014
    CONSOLE   = 0x00000000
    === boot complete ===

Reaching the completion breakpoint proved only that the transmitter ran; it
could not have caught either fault.  Comparing bytes is what did.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Added generic timer support to the S32Z280-594EVB example

The Arm generic timer now runs on silicon: CNTFRQ is programmed at EL2 and a
one-shot compare fires when it should.

    CNTFRQ programmed = 0x007A1200 (8000000)
    elapsed counts    = 8000043   (requested 8000000, +0.001%)
    host interval     = 1.0033 s

Both halves matter.  The elapsed count shows the counter and the compare agree
with each other; the host interval shows the rate is actually 8 MHz rather than
merely self-consistent, which a wrong CNTFRQ would also produce.  The residual
43 counts is about 5 us of polling overhead.

CNTFRQ reads zero out of reset here, exactly as on the Armv8-R AEM FVP -- so
that FVP behaviour is real silicon behaviour, not a model artefact, and the
port readme's warning to re-verify it was well placed.  The parallel stops
there: the FVP also leaves the system counter itself stopped and its BSP must
start it, whereas on this board the BootROM has the counter running and only
the software-declared constant was missing.  CNTFRQ is writable only at the
highest implemented exception level, so entry.S programs it before the drop to
EL1.

8 MHz was established three independent ways rather than assumed:

  - Measured: CNTPCT sampled against host wall-clock time over a 32-second
    interval gave 8.0227 MHz.
  - Derived: RTU.GPR CFG_CNTDV reads 4, so the divider is (4+1) = 5, and the
    board's FXOSC is 40 MHz -- itself already corroborated by the LINFlexD
    baud divisors the BootROM left behind.  40 / 5 = 8.
  - Confirmed: the one-shot compare above.

The compare is polled with the interrupt masked, deliberately.  There is no
GIC configured yet, and proving the timer counts and fires first means a later
interrupt failure cannot be confused with the timer itself being wrong -- the
same staging that made the console tractable.

Also recorded, from a detour that cost a run: RTU0.GPR CFG_CNTDV at 0x76120010
is readable by the debugger, which returns 4, but a load from the core at EL1
kills the image.  No fault handler runs, boot_stage stays at 4, and the debug
connection drops -- the signature of a stalled bus access rather than an abort,
and a failure mode that leaves nothing on-target to inspect.  The debugger
reaches that window over the AXI-AP, which does not go through whatever gates
the core's own access to RTU peripheral space.  The image no longer touches it;
the value lives in platform.h instead.  This is the second address on this part
that is debugger-visible but not core-reachable, so peripheral windows are now
worth probing from the core before relying on them.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Enabled the R52 peripheral port and located the GIC on S32Z280

Two peripheral windows on this part stalled the core outright when read: no
abort, no fault handler, boot_stage frozen at its last value, and the debug
connection dropping with it.  A stall gives you nothing on-target to inspect,
so both were found by printing a marker before each access and seeing which
marker came last.  The Cortex-R52 TRM r1p3 (100026_0103_00_en) explains both;
NXP's reference manual mentions neither.

1.  The low-latency peripheral port is disabled at reset.

    IMP_PERIPHPREGIONR (TRM 3.3.80) describes a region at 0x76000000 of 4 MB
    on this part -- the RTU peripheral space -- with separate enables for EL2
    and EL1/0.  Read from the board it was 0x76000034: base 0x76000000, size
    0b01101 = 4 MB, and both enable bits clear.  The TRM is explicit that each
    "resets to 0".  Until they are set, every access in that window stalls.

    entry.S now sets both at EL2, which is also where it has to happen: EL1
    writes to this register trap to EL2 when HACTLR.PERIPHPREGIONR is clear.
    PERIPHPRG now reads 0x76000037 and the core reads RTU0.GPR CFG_CNTDV = 4
    for itself -- the same value the debugger saw, which independently
    confirms the divider behind the 8 MHz counter rate from the core's own
    view rather than the debug path's.

2.  The GIC is where NXP says, and needs an MPU mapping to reach.

    IMP_CBAR (TRM 3.3.17) holds the physical base of the memory-mapped GIC
    distributor in bits [31:21], its reset value wired from CFGPERIPHBASE.
    Read from the board it is 0x47800000, confirming NXP's memory map from the
    hardware rather than from a vendor debugger script -- the reference manual
    contains no GIC base address anywhere.  NXP's cryptic note that the "R52
    Cluster set addr [31:21]" is just CFGPERIPHBASE[31:21] restated.

    TRM Table 9-1 gives the frame layout relative to that base: distributor at
    +0x000000, redistributor control at +0x100000, redistributor SGI/PPI at
    +0x110000, then +0x20000 per further core.  That matches what the FVP
    example's driver already assumes, so it should port with address changes.

    The distributor still stalls, and the TRM says why: "Ensure that the memory
    region used for the GIC Distributor is configured as Device nGnRnE."  The
    MPU is disabled here -- SCTLR.M reads 0 -- so nothing maps that region.
    Enabling the peripheral port does not help, the GIC being outside that
    window and reached over AXIM.  Reading it waits on MPU programming, which
    is the next piece of work; the probe is removed rather than left in, since
    while present it stalled every run and cost the console and timer results
    that precede it.

gic_probe.c holds the two system-register reads.  They cannot hang the bus,
unlike the memory accesses they diagnose, which is the point of them.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Reached the GIC on S32Z280 by mapping it Device nGnRnE

The GIC distributor and redistributor now respond on silicon:

    MPU rgns   = 0x00000003
    SCTLR      = 0x70C50839      (M set; caches still off)
    GICD_PIDR2 = 0x0000003B      architecture revision 3 = GICv3
    GICD_TYPER = 0x0248001E
    GICR_PIDR2 = 0x0000003B      redistributor at base + 0x100000

Three independent sources had to agree to get here, and NXP's reference manual
supplied none of them.  IMP_CBAR reports the distributor base from the hardware
(0x47800000).  Cortex-R52 TRM Table 9-1 gives the frame layout: distributor at
+0x000000, redistributor control at +0x100000, SGI/PPI at +0x110000, then
+0x20000 per further core -- the layout the FVP example's gicv3.c already
assumes.  And the TRM requires the region be Device nGnRnE, which was confirmed
rather than taken on faith: mapped Normal write-back the distributor stalls the
core outright, mapped Device it reads 0x3B.

mpu.c is ported from the FVP example, keeping its PRBAR/PRLAR encoding and the
reversed-AP-bit-order finding, with an S32Z280 region table.  Caches are left
off: the point of this pass was reaching the GIC, and the debugger writes this
image straight into SRAM behind the caches, so enabling C and I deserves its own
step and its own check.

Diagnosis needed two new pieces of machinery, both kept:

  - Memory-resident progress markers (probe_stage), mirroring the console
    markers.  Enabling the MPU can take the core down in a way that records no
    fault AND takes the console with it, so boot_stage alone could not say how
    far a risky sequence got.  With markers inside mpu_init the failure went
    from "somewhere in the MPU code" to "the SCTLR.M write itself" in one run.

  - Resolving symbol addresses from the current ELF with nm on every
    post-mortem.  Adding probe_stage to the .data block shifted fault_vector,
    so reusing addresses across builds reads the wrong words -- which nearly
    produced a conclusion from a stale read.

OPEN QUESTION, recorded in mpu.c rather than papered over.  A five-region map
that adds a second Device window for the RTU peripheral space (0x76000000, 4 MB)
makes the SCTLR.M write kill the core: all regions program without error,
probe_stage reaches 0x42, and the marker one instruction later never lands.  No
exception is taken, so it stalls rather than aborts.  Coverage is not the
explanation -- that map spanned the whole address space, and an earlier comment
of mine claiming otherwise was wrong and has been corrected.  The cause is not
known.  Consequence: RTU peripheral access after the MPU is enabled is untested;
the CFG_CNTDV read this image does happens before the enable.

Also not tightened yet: code is writable and data executable in this map.  Wrong
for a protection demo, right for bring-up where an unmapped address costs a
silent stall rather than a fault.  Narrowing it belongs with the ThreadX
integration, which can verify it.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Fixed an out-of-bounds MPU table write that made enabling the MPU stall the core

The five-region map now programs correctly and the MPU enables with the GIC
reachable:

    R  idx  PRBAR      PRLAR
    R  0    00000000   477FFFC1
    R  1    47800002   479FFFC3    GIC, XN, Device nGnRnE
    R  2    47A00000   75FFFFC1
    R  3    76000002   763FFFC3    RTU peripherals, XN, Device nGnRnE
    R  4    76400000   FFFFFFC1
    MPU rgns = 5, SCTLR = 0x70C50839 (M set)
    GICD_PIDR2 = 0x3B, GICD_TYPER = 0x0248001E, GICR_PIDR2 = 0x3B

The bug was mine and it was simple: mpu_regions[] is declared with a fixed
size, and the FVP example this file came from sized it [3] to match the three
regions it programs.  A five-region table wrote mpu_regions[3] and [4] past the
end of the array.

What made it hard to see is what the corruption produced.  Every region
appeared to program without error, and the failure was a stall on the SCTLR.M
write -- no abort, no fault handler, no exception, and the debug connection
dropping.  Reading the regions back out of the hardware, with the MPU
deliberately left disabled so the readback could not itself stall, showed
region 3 with base 0x00000000 instead of 0x76000000.  With its correct limit
that region spanned 0x00000000-0x763FFFFF and overlapped regions 0 to 2, and
PMSAv8-R leaves overlapping regions UNPREDICTABLE -- so stalling was permitted
behaviour rather than a hardware fault.

One bug accounts for every observation: three regions worked, five failed, and
it failed identically whether the extra window was Device or Normal, because
the memory type was never involved.

Fixes: the table capacity is now named (MPU_TABLE_REGIONS) and sized 16, and
mpu_init checks mpu_regions_used against it as well as against MPUIR.  The
existing check only compared against the hardware's region count, which said 20
and told us nothing about the array.

Two earlier claims of mine, corrected here rather than left standing.  A comment
asserting "narrow coverage was the problem" was wrong: the failing map covered
the whole address space.  And this is NOT a defect inherited from the FVP
example -- that code declares [3] and uses three regions, entirely
self-consistent.  Worth offering upstream only as hardening: the bare [3] with
no bound check is what let a growing map overrun it silently.

The region readback is kept in the image.  It costs eight lines a boot and
verifies the map against the hardware every time, which is precisely what
turned an unexplained stall into a one-line diagnosis.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Enabled interrupt-driven ticks by clearing SCTLR.TE

Timer interrupts are delivered, acknowledged and re-armed on silicon:

    SCTLR       = 0x30C50838 -> 0x30C50839   (TE cleared, M set)
    IRQ count   = 7
    timer INTID = 30                          measured, not assumed
    spurious    = 0
    unexpected  = 0

SCTLR.TE -- Thumb Exception enable -- RESETS SET on this part.  Every vector
table in entry.S is A32, so the core entered each exception in T32 state,
decoded an A32 branch as Thumb, landed mid-instruction at a misaligned address
and took an undefined-instruction exception, which repeated the same way.  No
handler ever ran.  entry.S now clears it at EL1 and clears HSCTLR.TE at EL2.

The consequence was worse than a crash: it made every fault invisible.
boot_stage stayed at its last value, fault_vector stayed zero, and the core sat
at el1_vectors+4.  A data abort from an unmapped peripheral, an abort from an
overlapping MPU region, and a correctly delivered timer interrupt all presented
identically -- an unexplained hang with nothing recorded.  What gave it away was
reading the banked registers: LR_irq and SPSR_irq showed an IRQ had been taken
from SVC with interrupts enabled, CPSR showed UND mode with T set, and LR_und
pointed at a misaligned address inside the A32 vector table.

The Armv8-R AEM FVP resets with TE clear, which is why the same vector tables
work there and why nothing in the FVP work anticipated this.

This commit also corrects earlier comments of mine.  Several places described
these failures as "a stalled bus access rather than an abort", with the absence
of a recorded fault as the evidence.  That was wrong: they were ordinary
exceptions whose handlers were unreachable.  The two underlying bugs -- the
peripheral port disabled at reset, and the out-of-bounds MPU table write -- were
real and are correctly fixed, but the explanation for why they presented so
mutely was not.  Fixed in platform.h, mpu.c and bsp_boot.c.

Interrupt plumbing: gicv3.c is ported from the FVP example with the frame bases
in platform.h, derived from IMP_CBAR and TRM Table 9-1.  el1_irq_entry in
entry.S uses the classic A32 form rather than ThreadX's context save/restore,
since this image does not link the kernel.  irq_dispatch.c counts ticks instead
of calling _tx_timer_interrupt, and keeps separate spurious and unexpected-INTID
counters so that "no interrupt arrived" and "an interrupt arrived and was
mishandled" cannot be confused.  timer_start_oneshot_irq arms with IMASK clear;
the polled timer_start_oneshot keeps IMASK set, and the two are separate
functions because either mistake is silent.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Ran ThreadX on S32Z280 silicon with threads, tick and preemption

The kernel runs on the board:

    === ThreadX Cortex-R52 :: S32Z280-594EVB ===
    console   = 0x00000000
    entering kernel

    === ThreadX on S32Z280: results ===
    ticks     = 0x00000064      100 ticks
    sleeper   = 0x00000014       20 wakeups
    spinner   = 0x00159B48    1415496 iterations
    preempt   = 0x00000014       20 preemptions
    PASS threads, tick and preemption all verified

The figures are internally consistent, which is what makes them evidence rather
than encouragement: 100 ticks elapsed for a 100-tick sleep, so the tick rate is
exactly as configured; 20 sleeper wakeups is precisely 100/5 for a 5-tick sleep;
and 20 preemptions means every wakeup took the core from a runnable
lower-priority thread rather than receiving it by cooperative handoff.

This validates the context-switch assembly merged in #579 on real Cortex-R52
silicon.  Until now it had only run against the Armv8-R AEM FVP, which reports
MIDR part 0xD0F -- an architecture envelope model, not an R52 implementation.

The demo judges four things separately because on first silicon they fail
independently: that time advanced (the GIC delivers the timer PPI and
_tx_timer_interrupt is reached), that tx_thread_sleep returned (tick-driven
scheduling, not merely the interrupt), that the lowest-priority thread ran (the
sleeper actually yielded), and that the sleeper resumed while the spinner was
runnable (preemption).  A demo that only counted whether both threads ran would
pass with a broken tick if they happened to yield to each other.

Integration pieces, all from the FVP example with the board's differences:

  - tx_initialize_low_level.S publishes the system stack and the first free
    address, then calls board_init().  Shorter than the A-profile reference
    ports because entry.S has already given every mode its own stack from
    dedicated linker regions.

  - entry.S routes the IRQ vector through _tx_thread_context_save and
    _tx_thread_context_restore under TX_R52_USE_THREADX_IRQ, keeping the
    standalone A32 handler for s32z280_boot.elf, which does not link the kernel
    so that a boot failure there stays unambiguous.

  - board_init() runs mpu_init() FIRST.  The GIC distributor is unreachable
    until its region is mapped Device nGnRnE, so any GIC access before the MPU
    is enabled aborts.

  - link.lds provides _end as well as end; _tx_initialize_low_level publishes
    unused memory from _end.

  - bsp_done() is defined in the demo rather than shared from bsp_boot.c, whose
    bsp_main would collide with the demo's.

s32z280_boot.elf still builds and the FVP example still passes 6/6.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Tightened the S32Z280 MPU map and verified protection is enforced

The map now carries real permissions -- code read-only and executable, data
writable and never executable, peripherals Device nGnRnE and never executable,
everything else unmapped -- and the enforcement is demonstrated rather than
asserted:

    X1 write to read-only code region
    faults     = 0x00000001
    fault DFSR = 0x00000A0C      bit 11 (WnR) set: a write fault
    fault addr = 0x79900000      the address written
    X2 PASS write to code faulted and was recovered

Exactly one fault, identified as a write, at the precise address, followed by
successful resumption.  ThreadX still passes on the same map: ticks 100,
sleeper 20, spinner ~1e6, preempt 20, and IRQ count keeps advancing during the
protection test.

Reading SCTLR back or listing the programmed regions would only show what was
configured.  Provoking the violation shows it is enforced, which is the claim
that matters for a protection story.

entry.S gains a recoverable data-abort path, gated on a fault_expected flag the
test arms.  It records DFSR and DFAR, counts the fault, and resumes at the
instruction after the faulting access -- lr on data-abort entry is the faulting
address plus 8, so subs pc, lr, #4 skips the access instead of retrying it
forever.  Unarmed, a data abort stays fatal and reported, which is what a real
bug should get.

This supersedes the permissive bring-up map, and the reason that map existed is
worth recording: a narrow map had failed earlier and the cause was unknown, so
coverage was blamed.  Coverage was never the problem.  An out-of-bounds write to
a [3]-sized region table put a region at base 0 overlapping everything, and
SCTLR.TE being set meant the resulting abort could not reach a handler.  With
both fixed a narrow map works first time, and -- more usefully -- a mistake in it
now reports vector, syndrome and faulting address instead of hanging mutely.

Until TE was cleared no fault handler on this board could run, so no protection
claim about it could be tested at all.  This is the first one that could.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Enabled the S32Z280 caches, with an honest effectiveness result

Both caches are on and verified from SCTLR:

    CLIDR      = 0x09200003
    SCTLR      = 0x30C50839 -> 0x30C5183D    (C and I set)
    cachesOn   = 1
    counts off = 0x00094DF4
    counts on  = 0x00094CC8
    gain/1000  = 0
    C3 PASS caches enabled (SCTLR.C and SCTLR.I set)
    C4 no significant speedup -- expected here, the workload is already
       in fast local SRAM

cache.c supplies what the FVP example explicitly deferred to silicon: a
data-cache set/way sweep that reads CLIDR and CCSIDR and walks every set and
way the hardware reports, rather than assuming a geometry.  The FVP file notes
that this "belongs with the silicon bring-up where it can be verified against
the real cache geometry"; this is that.  The sweep matters because caches are
not architecturally guaranteed invalid out of reset, and enabling a write-back
data cache holding stale valid lines would evict them over live memory.

The effectiveness result is reported separately from the enable, and that
separation was earned.  The first version of this test passed on warm < cold and
printed "caches on and the workload got faster" for 609904 counts against
609686 -- a 0.036% difference that is measurement noise.  That criterion was
worthless and the PASS was misleading.  It now reports the gain as a fraction
and only claims a speedup above 10%.

No speedup here is a plausible result rather than a defect: both the code and
the data this workload touches live in RTU-local low-latency SRAM, close to core
speed already.  There is no slow memory on this board to demonstrate a cache
against -- DDR would be the place, and DDR is not initialised.  So the honest
claim is "enabled and harmless", not "enabled and beneficial".

cache_clean_all() is called before parking, and this is not optional.  With a
write-back data cache, values this image writes can sit dirty in cache where a
debugger reading SRAM cannot see them -- and the memory-resident progress
markers and fault records this bring-up depends on are read exactly that way.
Without the clean, a post-mortem read could report stale values and look like a
fault that never happened.

ThreadX still passes on the same configuration and the protection test still
faults correctly: write to the read-only code region takes exactly one write
fault at the written address and recovers.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Made the cache benchmark actually measure the cache

The caches now show a real effect instead of noise:

    D line     = 64 bytes
    D ways     = 4
    D sets     = 64
    D bytes    = 0x4000    16 KB L1 data cache, read from CCSIDR
    bench wds  = 0x800     8 KB working set, half the cache
    counts off = 0x000D1B87   858503
    counts on  = 0x0009F515   652053
    gain/1000  = 0xF0         24% faster
    C4 PASS caches measurably faster (>=10%)

The previous benchmark could not have measured anything, and its 0.036%
"speedup" was noise.  Two things were wrong with it.

It used a static buffer in the RTU-local fast-data bank.  That memory is close
to core speed already, so caching it saves almost nothing -- there is no latency
to hide.  The benchmark now runs over the extended SRAM at S32Z_EXT_SRAM_BASE,
which the Reference Manual describes as NOT RTU-local, and mpu.c gains a sixth
region to map it Normal write-back and never executable.

And its working set was a guessed 4 KB constant.  It is now sized at run time
from CCSIDR to half the reported data cache, so it fits and is revisited every
pass.  A set larger than the cache would stream through and evict, showing
little benefit even over slow memory -- a guess could have produced a null
result for a reason that has nothing to do with whether the cache works.

24% on this loop is a believable figure rather than a suspiciously large one:
the workload is a store-heavy read-modify-write over a write-back cache, so part
of the cost is write traffic that still has to reach SRAM.

Useful by-product: the L1 data cache geometry is now read and printed rather
than assumed.  Nothing in the SoC reference manual gives it, and the set/way
invalidate sweep in cache.c depends on it being right.

ThreadX still passes and the protection test still faults correctly.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Fixed the PRBAR field encoding and proved execute-never enforcement

program_region() shifted every PRBAR field one bit too far left: SH<<4,
AP<<2 and XN<<1, where the Cortex-R52 TRM (r1p3 figure 3-39, table 3-80)
places BASE[31:6], RES0[5], SH[4:3], AP[2:1] and XN[0].  The intended XN
therefore landed in AP[1] -- the EL0-access bit -- and no region was ever
execute-never.  On this board the core ran instructions straight out of
.data with LR_abt at 0x3178000c to prove it.

The write-to-read-only test passed throughout, which is why this survived:
under the old shift the AP value's low bit happened to land in the real
AP[2], the read-only bit, so permissions came out right by accident while
XN was silently discarded.  Only an instruction fetch from a region marked
non-executable could distinguish the two.

The AP macros in mpu.h go back to the architectural encoding, AP[2] for
read-only and AP[1] for EL0 access.  No calibration is needed; the values
were correct as published all along.

Verified on S32Z280 silicon: the data region now reads back PRBAR
0x31780001 with XN set, a write to the read-only code region faults with
DFSR 0xA0C, execution from the data region takes a prefetch abort with
IFSR 0x20C at 0x31780000, both faults recover, and the ThreadX demo still
reports 100 ticks with 20 preemptions.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Corrected the same PRBAR encoding bug in the FVP MPU example

The FVP example carries the identical off-by-one shift that the S32Z280
bring-up exposed, so no region there was execute-never either: the data
region and the whole peripheral region were both freely executable, and
each also gained unintended EL0 access from the misplaced XN bit.

This also retires the "PRBAR.AP bit order is reversed" claim in mpu.h.
That note rested on a real measurement -- four regions, one per AP
encoding, privileged write attempted on each, writes faulting for 0b01 and
0b11 -- but the cause was the shift, not the bit order.  With AP written
into bits[3:2], its low bit lands in the real AP[2], the read-only bit, so
writes fault exactly when that bit is set.  The "high bit grants EL0
access" half of the conclusion was never tested; under the old shift that
bit landed in SH[0], programming a shareability the TRM calls
UNPREDICTABLE.  The architectural encoding needs no calibration.

demo_mpu.c gains the check that would have caught this: an instruction
fetch from the execute-never data region must take a prefetch abort.
Recovering from one cannot work the way the data-abort path does, since
skipping the faulting instruction is impossible when the instruction is
what could not be fetched, so entry.S gains a recoverable prefetch-abort
path that returns to a landing point recorded by mpu_try_execute.

Verified both ways.  All five FVP images pass, and the new check reports
prefetch aborts 0 -> 1.  Restoring the old shift for one run makes it fail
with 0 -> 0, so the check is not vacuous.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Removed the unassemblable .arch directive from the S32Z280 example

GNU as accepts ".arch armv8-r"; LLVM's integrated assembler accepts no spelling
of it -- not armv8-r, not armv8r, not armv8-r+crc -- and stops with "Unknown
Arch: armv8-r". There is nothing to substitute, so the directive goes. The
architecture comes from -mcpu=cortex-r52 on the command line, which every
toolchain file passes, so it only restated it.

entry.S in this example never had one. The two equivalents in the FVP example are
handled separately, in the change that brings those images under the LLVM check.

Verified: both S32Z280 assembly sources now assemble with Arm Toolchain for
Embedded 22.1.0, and the GNU build of s32z280_boot.elf and s32z280_demo.elf is
unchanged with no warnings.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Stopped read_identity.gdb reporting failure after a successful demo run

The script reads r52_identity, which only the boot image defines. The kernel demo
shares the same bsp_done breakpoint and the same boot_stage marker, so it is
convenient to point the same script at either image -- but against the demo the
lookup raised and gdb exited 1, after the demo had already printed
"PASS threads, tick and preemption all verified" over the console. A tool that
reports failure on success is worse than one that prints nothing.

The lookup is now guarded, and says which image it is looking at rather than
falling over. Everything after it is skipped when the structure is absent.

Verified against both images on S32Z280 silicon: the demo exits 0 and reports its
own PASS, and the boot image still prints the full identity report with MIDR
0x411FD133 and C3, C4, X2 and X4 all passing.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-12 10:25:35 -04:00
Frédéric Desbiens b0ad67bae9 Brought the Cortex-R52 examples under the LLVM check, and fixed two blockers (#604)
#600 added a line reporting which example builds the LLVM check passes over, and
cortex_r52 was on it: its examples are driven by CMake rather than by a
build_threadx.sh pair, so the linking stage never touched them. Covering them
turned up two reasons they could not have been built with anything but GNU.

.arch armv8-r has no portable spelling. GNU as accepts it, and LLVM's integrated
assembler rejects every variant -- armv8-r, armv8r, armv8-r+crc -- with "Unknown
Arch: armv8-r". There is nothing to substitute, so the directive is gone from
entry.S and tx_initialize_low_level.S; -mcpu=cortex-r52 already selects the
architecture, both toolchain files pass it, and the directive only restated it.
Worth noting where this hid: the assembly stage walks ports/*/gnu/src, so example
assembly had never been assembled by LLVM at all.

-Wl,--no-warn-rwx-segments is GNU ld only, added in binutils 2.39. ld.lld does
not warn about RWX segments and rejects the flag outright, failing the link with
"unknown argument". It is now selected on CMAKE_C_COMPILER_ID rather than spelled
into all six targets, and the reason it exists at all -- a bare-metal image has
one flat DRAM region and leaves access control to the MPU -- moves to the one
place that sets it.

cmake/cortex_r52_clang.cmake is the toolchain file. It names the tools as found
on PATH, which is what CI uses, then pins $HOME/toolchains if that directory
exists, mirroring how cortex_r52.cmake pins the GNU toolchain and for the same
reason. Falling back rather than requiring the pinned path keeps the file usable
on a machine that keeps clang elsewhere. THREADX_TOOLCHAIN stays "gnu": there is
no clang port directory, this builds the gnu sources with a different compiler,
which is what the whole check does for every other Arm port.

check_clang.sh gains a fourth stage for CMake-driven examples, and no longer
reports cortex_r52 as a gap. It reads the image list out of the generated ninja
graph rather than repeating it, so adding a target cannot escape the check, and
filters out the cmake_object_order_depends_target_* phonies -- counting those
reported ten images where there are five.

Verified with Arm Toolchain for Embedded 22.1.0, the version the workflow pins.
All five images link, and all five then run and pass on FVP_BaseR_AEMv8R:
boot_check, demo_m2, demo_m3, demo_threadx and demo_mpu. That is a step beyond
the AArch64 examples, which are link-verified only. The full check reports 711 of
711 assembly sources, 185 of 185 common C sources for each of nine cores, 42 of 42
script-driven examples and 5 of 5 CMake images, leaving only cortex_a5_smp,
cortex_a7_smp and cortex_a9_smp listed as having no driver.

GNU is unaffected: the same five images build with no warnings and the FVP test
suite passes 5 of 5.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-11 19:53:11 -04:00
Frédéric Desbiens 31fbe21d20 Removed developer machine paths from the shipped project files (#603)
The Arm Development Studio launch configurations and the IAR project files
carried absolute paths from the machines they were last opened on: four
individuals' home directories, a Broadcom support build tree, and directories
such as C:\release\threadx, C:\temp1702\tx and D:\threadx. They are published in
the repository and none of them resolves for anyone else.

In the launch files all of it is saved session state, which files that were never
opened in a debugger simply do not have:

  breakpoints                  72 files, saved breakpoints anchored to source
                               paths under cortex_a35 on one machine. Emptied
                               rather than deleted, matching the 27 launch files
                               that already carry an empty list.
  scripts_view_script_links    70 files, a scripts view cache. The script the
                               launch actually runs is named one line earlier
                               through ${workspace_loc:...}, which is portable.
  substitutePath               65 files, source lookup remapping. 30 map a path
                               to itself, and the rest point at one machine's
                               copy of libgloss or a build directory.
  TREE_NODE_PROPERTIES         39 files, which rows the variables view had
                               expanded.
  DebugCommandLine.History      2 files, the debug console command history,
                               including a typed absolute path.

The two IAR cases are not session state and are corrected rather than removed:

  IarchiveOutput               23 files. The librarian output path pointed at
                               another machine, so the library was written
                               outside the project. Set to ###Unitialized###,
                               which 86 other option blocks already use, letting
                               IAR derive $EXE_DIR$\$PROJ_FNAME$.a.
  IlinkIcfFile                  1 file. The linker configuration file is load
                               bearing, so the file name is kept and the path
                               made project relative, as 33 other option blocks
                               spell it.

Seven of the files are ports_arch templates, so the generators would otherwise
have copied the paths back.

Verified: all 162 launch and IAR project files in the tree parse as well-formed
XML afterwards, no tracked launch or project file mentions any of the four user
names, and scripts/check_ports.sh passes, so the templates and their generated
copies still agree.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-11 19:39:41 -04:00
Frédéric Desbiens c78c080139 Made the ARMv8-A debug launch configurations name their own core (#602)
Two identifiers in the Arm Development Studio launch configurations still said
A35 in every AArch64 port, because the generator's patch table could not reach
them.

The config database taxonomy id is spelt in lower case in the launch files,
/platform/armfvp/base_a35x4, but the search string read base_A35x4. Both
generators are case sensitive here, sed in update.sh and -creplace in
update.ps1, so the rule matched nothing and all 27 ports kept the A35 taxonomy
id while the platform name and model beside it named the right core. That the
rule exists at all shows the substitution was intended, and the ARMv7-A
generator does the same thing correctly, which is why its ports carry a per-core
ve_cortex_a<n>x1 id.

The SMP launch files name the activity "Cortex-A35x4 SMP" where the ThreadX ones
say "Debug Cortex-A35", so the existing activity rule reached the ThreadX ports
only and every SMP port advertised an A35 activity. Add a rule that matches the
" SMP" suffix, which also keeps it from rewriting the FVP model name that the
line above already handles.

Both changes are mirrored in update.ps1, which stays equivalent to update.sh.

Verified by regenerating: 50 launch files change and nothing else, 100 lines of
taxonomy id across 24 ThreadX and 26 SMP files, and 52 lines of activity name
across the 26 SMP files. No other attribute in those files moves, so the new
" SMP" rule does not touch the model name. The two Cortex-A35 ports are
untouched, as they must be, since substituting A35 for A35 is a no-op.
scripts/check_ports.sh passes, so the result is reproducible.

This cannot be exercised here: confirming that Arm Development Studio resolves
the per-core taxonomy ids needs Arm Development Studio. The change rests on the
platform name and model already naming the core, on the patch rule's existence,
and on the ARMv7-A ports having shipped per-core ids all along.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-11 16:20:48 -04:00
Frédéric Desbiens 96b1638fe4 Gave the Armv8-M examples a linkable sample and fixed what stopped them (#601)
The Cortex-M23, M33, M55 and M85 examples each shipped a build_threadx.sh with
no build_threadx_sample.sh, so scripts/check_clang.sh skipped all four at its
linking stage: the whole Armv8-M family had no example-link coverage. Adding the
missing script surfaced five separate reasons why none of them could have linked.

ports/cortex_m23/gnu/src/tx_initialize_low_level.S referenced
Image$$ARM_LIB_STACK$$ZI$$Limit and __Vectors, which are Arm toolchain
scatter-load names that no GNU linker defines. This is a GNU port, so it now
uses __RAM_segment_used_end__ and _vectors, as the Cortex-M33 version already
does. The Cortex-A ports mention Image$$ZI$$Limit only inside a comment; this
was a live relocation.

The Cortex-M23 crt0 needed unified assembler syntax. The Armv7-M file it derives
from carries no .syntax directive, so GNU as reads it in the legacy divided
syntax, where a Thumb data-processing instruction sets the flags whether or not
the mnemonic says so, and "mov r2, #0" assembles to 2200, MOVS. LLVM implements
unified syntax only, where those mnemonics mean the non-flag-setting forms that
Armv8-M Baseline does not have. Disassembling the Cortex-M4 object confirms GNU
already emits 2200 movs, 1a52 subs and 3001 adds, so spelling them out changes
no encoding. The flags carry meaning: crt0_memory_copy branches on the result of
"subs r2, r2, r1" and again on "subs r2, #1".

Cortex-M23 has no SVC_Handler to name in its vector table. Its library is built
-DTX_SINGLE_MODE_NON_SECURE, and tx_thread_schedule.S defines that handler only
when neither TX_SINGLE_MODE_SECURE nor TX_SINGLE_MODE_NON_SECURE is set, so the
SVCall slot takes __tx_BadHandler. The table also uses the CMSIS handler names
throughout, because that is what the Armv8-M ports export: the Armv7-M table
references __tx_SVCallHandler and __tx_SysTickHandler, which resolve to nothing
here.

The other three needed a C library. tx_thread_secure_stack.c calls malloc and
free, which pulls the allocator in, and the sample scripts already defined
SYSCALL_LIB for exactly that without ever passing it to the linker. Their linker
scripts now also provide end and _end, which newlib's libnosys _sbrk wants, and
__heap_start and __heap_end, which picolibc wants instead.

Cortex-M55 and M85 build with the hard-float ABI now, in both the library and
the sample. Arm Toolchain for Embedded ships no soft-float MVE multilib and says
so plainly: "No library available for MVE with soft-float ABI." The library and
the sample have to agree, and -mfloat-abi=hard is what check_clang.sh already
uses for these two cores in PORT_TARGET.

Verified with GNU 13.2.1 and Arm Toolchain for Embedded 22.1.0. Every port links
under both: Cortex-M23 at 206,136 and 312,636 bytes, M33 at 305,856 and 320,696,
M55 at 306,228 and 320,664, M85 at 306,240 and 320,668. check_clang.sh now
reports 42 of 42 example builds linking, up from 38, and no longer lists any port
as carrying build_threadx.sh without build_threadx_sample.sh. The images are
link-verified only and have not been executed, as with the AArch64 examples.

Only .sh drivers are added. Cortex-M23 has a build_threadx.bat with no sample
counterpart and the other three have no .bat at all; adding untested Windows
scripts belongs in its own change.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-11 14:00:54 -04:00
Frédéric Desbiens b87a62d218 Brought Cortex-A34 and the A78 SMP port under the ARMv8-A generator (#599)
ports/cortex_a34, ports_smp/cortex_a34_smp and ports_smp/cortex_a78_smp were
absent from the ARMv8-A generator's core list, so ports_arch never reached them
and scripts/check_ports.sh could not detect that they had drifted. They had.

The two Cortex-A34 ports sat two releases behind the shared sources, at 6.1.10
against 6.3.0. The difference is not only banners: tx_initialize_low_level.S
lacked the SUB x1, x1, #15 that precedes the BIC when the system stack pointer
is recorded, so the value stored was the incoming SP rather than the first
16-byte boundary below it. The Arm Development Studio example was missing
GICv3_aliases.h, which every other port's GICv3_gicc.h includes to reach the
interrupt controller registers through the stringify indirection.

The Cortex-A78 SMP port had never had its debug launch configuration patched at
all. It still named the A35 platform and model, so Debug Cortex-A35,
Base_A35x4 and FVP_Base_Cortex-A35x4 would have started an A35 model for an A78
port.

Add cortex_a34 to the core list. Cortex-A78 cannot go in the same list, because
the generator walks cores against port sets and would then create a
ports/cortex_a78 that the tree has never had; give it a separate SMP-only list
consulted only for the tx_smp port set. Both changes are mirrored in update.ps1,
which stays equivalent to update.sh.

Verified by regenerating: exactly these three port directories change, 116 files
modified and 13 added, and no already-covered core moves, so the core list was
the only thing holding them back. scripts/check_ports.sh passes, which now means
these three are checked for reproducibility for the first time.
scripts/check_clang.sh reports 38 of 38 example builds linked, the three new
ports included.

readme_threadx.txt is not added here. Only the two Cortex-A35 template ports
carry it; the other 23 AArch64 ports do not, so its absence from Cortex-A34 is
the tree's norm rather than drift, and supplying it everywhere is separate work.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-11 09:57:55 -04:00
Frédéric Desbiens 51ca412b36 Added build scripts for the AArch64 examples, which had none (#597)
The AArch64 examples could only be built through the Arm Development Studio
project files beside them. There was no script, so nothing in CI or on a
developer's machine could build them, and the sources they need were never
named anywhere: vectors.S, v8_aarch64.S and v8_utils.S were present but
unreferenced, which is why a hand written link failed on GetCPUID, GetAffinity,
InvalidateUDCaches and ZeroBlock. Those four are defined in v8_aarch64.S and
v8_utils.S, in the tree all along.

Add one pair of scripts, in ports_arch so every AArch64 port receives them.
Both derive the -mcpu value and the kernel source directory from the port
directory they sit in, so a single pair serves the twelve ThreadX ports and the
twelve SMP ports, the latter building against common_smp and picking up the C
sources those ports carry alongside the assembly. TOOLCHAIN=atfe selects Arm
Toolchain for Embedded in place of the GNU toolchain, as it does for the ARMv7
example scripts.

One symbol genuinely had no definition. startup.S calls
initialise_monitor_handles to open the standard file handles over a debugger
connection; the GNU toolchain provides it in libgloss through
--specs=rdimon.specs, while picolibc has no equivalent and neither does the
LLVM toolchain's semihosting library. semihost_stub.S supplies a weak no-op,
linked for that toolchain only, so a real definition always wins.

Extend scripts/check_clang.sh to cover ports_smp as well as ports, since the
example builds now exist there too.

Verified by building every AArch64 example with Arm Toolchain for Embedded:
twelve ThreadX ports at 314,744 to 315,256 bytes of text and twelve SMP ports
at 325,304 to 325,880. scripts/check_clang.sh now reports 35 of 35 example
builds linking with none failing, the twenty-four new ones plus the eleven that
already built. The images have not been executed: the sample
targets the Base platform peripheral addresses, which the available emulator
does not provide.

Cortex-A34, and the Cortex-A34 and A78 SMP ports, are left out. They are not in
the generator's core list, so they receive nothing from ports_arch and cannot be
checked for reproducibility either. Bringing them in rewrites 114 files in those
three directories, which deserves its own change.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-11 08:54:04 -04:00
Frédéric Desbiens 08eff8061c Made the Cortex-A12, A15 and A17 examples link with an LLVM toolchain (#596)
Those three sample scripts compiled crt0.S and reset.S but named neither on the
link line, and omitted -nostartfiles, so the toolchain was also free to link its
own startup. Under GNU that produced a working image; under an LLVM toolchain
picolibc's crt0 was pulled in and wanted __data_start, __data_source,
__data_size, __bss_size, __stack, __tls_base and __arm32_tls_tcb_offset, none of
which the linker script defines.

Defining picolibc's contract in the shared linker script was tried first and
abandoned: it reaches into ABI constants that cannot be verified here, and it is
unnecessary, because the in-tree crt0.S is already a complete startup for this
script. It sets up the stack and zeroes BSS between __bss_start__ and
__bss_end__, and .data carries no AT() so there is nothing to copy.

So the three scripts now pass -nostartfiles and name crt0.o and reset.o, which
is what the four sibling A profile cores already do and what these scripts were
evidently compiling those files for.

Both the old and the new GNU images contain the in-tree startup, __vectors from
reset.S and _mainCRTStartup from crt0.S, so the startup was already being used
through implicit startup file resolution rather than an explicit operand. How
that resolution happened without the objects being named is not accounted for
here, which is itself the argument for naming them: the same implicit behaviour
does not hold across toolchains, and that is why the LLVM link failed.

Add a readme to each of the three example directories. Their
tx_initialize_low_level.S is the generic ARMv7-A skeleton, 300 lines and
identical across all three, with no interrupt controller programming, no timer
and no vector table installation, where the A5, A7, A8 and A9 examples have all
three and ship the matching Versatile Express support files. These three
therefore demonstrate that the port builds; they will not receive a timer tick.
Nothing in the tree said so, which invites the assumption that they are
equivalent.

Verified with both toolchains for all three cores. GNU links at 41,276 bytes of
text, down from 41,768 because the unused toolchain startup is no longer
included, and Arm Toolchain for Embedded links at 37,982 where it previously
could not link at all. The resulting image is structurally sound: entry at
_start, __vectors at address zero, and a BSS range in RAM. It has not been
executed. scripts/check_clang.sh now links eleven of eleven example builds.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-10 17:41:23 -04:00
Frédéric Desbiens a8aebc5206 Completed the Cortex-M0 barriers and made its example build with both toolchains (#595)
Two unrelated Cortex-M0 gaps, both left over from earlier work.

The memory barriers from #523 reached the Cortex-M0 gnu port but not its ac6 or
iar siblings, in either the inline system return in tx_port.h or the assembly
routine. Both take the same GNU or IAR code path, and iar already carried the
entry barrier, so the missing pieces were the entry pair for ac6 and the barrier
after restoring the interrupt posture for both. All three tools now match.

The Cortex-M0 example could not link with any toolchain. cortexm0_crt0.S
references 23 linker script symbols and the script defined only 12 of them, so
__text_start__, __text_end__, __text_load_start__, the rodata and fast section
symbols, and the ctors and dtors load addresses were all unresolved. The
Cortex-M4 script defines all 23, including a .fast section with no content whose
symbols exist so that the startup copy is a no-op, and its comment says as much.
The Cortex-M0 script is brought to that same shape.

That left the example failing under LLVM only, on instructions that ARMv6-M can
encode just one way. The file declared .code 16 but no syntax mode, so GNU as
used the legacy divided syntax in which a plain add or sub sets the flags
implicitly, while LLVM implements unified syntax only and rejected the
non-flag-setting spelling. Declaring .syntax unified and writing movs, adds and
subs makes both assemblers agree, and the encodings GNU produces are byte
identical before and after, verified by disassembling both objects.

The Cortex-M0 example now links with GNU at 22,520 bytes of text and with Arm
Toolchain for Embedded at 22,866, so it comes off the list of examples not
expected to link and scripts/check_clang.sh now links eight of eight.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-10 17:34:13 -04:00
Frédéric Desbiens f3ab36dacc Made the example builds work with LLVM, and fixed four that were broken on Linux (#594)
* Made the example builds work with LLVM, and fixed four that were broken on Linux

The gnu example builds are the natural LLVM path: Arm Toolchain for Embedded is
LLVM based, consumes GNU ld linker scripts, and needs none of the scatter files
or Arm DS projects the ac6 examples carry. So rather than port the ac6 examples,
teach the gnu scripts to drive either toolchain.

Each script now selects the toolchain from a TOOLCHAIN variable that defaults to
gnu, so every existing invocation behaves exactly as before. TOOLCHAIN=atfe
switches the compiler, adds the target triple, passes the entry symbol through
to the linker rather than to the driver, and links the toolchain's own
semihosting library in place of --specs=nosys.specs, which is a GCC spec file
mechanism with no LLVM equivalent. That last choice avoids adding a syscall stub
source to every example.

The scripts are parameterised rather than duplicated. Copies of build scripts
would drift the first time one side was edited, which is the failure this
repository has just spent several changes recovering from.

Four sample scripts could not run on a case-sensitive filesystem at all. They
compiled MP_PrivateTimer.s and V7.s while the files on disk are
MP_PrivateTimer.S and v7.s, and the Cortex-A5 and A9 scripts named MP_GIC.s
where the file is MP_GIC.S. The link lines named V7.o accordingly. Corrected to
match the files, which is why the Cortex-A5, A7, A8 and A9 examples now build on
Linux where before they could not.

Extend scripts/check_clang.sh to link the example builds as well as compile the
sources, since compiling proves the sources parse while only linking exercises
entry symbols, linker scripts and the C library together.

Eight examples are listed as not expected to link, each with its reason, so the
gaps stay visible rather than being silently skipped. Five of them fail with the
GNU toolchain too and are therefore not LLVM problems: the Cortex-M0 example's
crt0 references __text_load_start__, __text_start__ and __text_end__, which its
linker script never defines, while the Cortex-M4 script defines the equivalents;
the arm9, arm11, Cortex-R4 and Cortex-R5 examples need newlib multilib variants
that are not present in every GNU toolchain packaging. The Cortex-A12, A15 and
A17 examples fail only with LLVM, because their link line omits -nostartfiles so
the toolchain's own crt0 is linked and wants picolibc's __data_start,
__data_source, __data_size and __bss_size, which their linker script does not
define.

Verified by building every example with both toolchains. Seven link with both:
Cortex-A5, A7, A8, A9, M3, M4 and M7. The GNU results are unchanged where they
worked before, and now also succeed for the four scripts with the case bug.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Resolved the compiler path before running the example builds

The example stage runs each build script from inside its own directory, so a
relative path passed with --clang stopped resolving there and every example
build failed instantly. Local runs passed an absolute path and did not show it;
the CI job passes a path relative to the workspace root, which did.

Resolve the compiler to an absolute path once, before any directory change, and
print it so the toolchain in use is visible in the log.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Parameterised the archiver as well, and stopped the check hiding failures

The toolchain selection covered the compiler but not arm-none-eabi-ar, which the
scripts invoke 628 times to assemble the library archive. A machine with an LLVM
toolchain and no GNU one therefore could not build any example, which is exactly
the situation in CI: the runner has no Arm GNU toolchain installed. Local runs
had one, so this only appeared once the job ran.

Select the archiver alongside the compiler, taking llvm-ar from beside clang in
the toolchain. The four scripts that call arm-none-eabi-ld directly are left
alone: they are arm9, arm11, Cortex-R4 and Cortex-R5, all already listed as not
expected to link, and their invocations are specific to GNU ld in ways that
parameterising would not resolve.

The check reported "example build produced no image" and then filtered the log
for lines containing "error", which hid the actual cause, since a missing tool
reports "command not found" or "No such file or directory". It now prints the
tail of the log. That filtering cost two CI round trips to diagnose something
the first run already knew.

Verified by shadowing arm-none-eabi-gcc, arm-none-eabi-ar, arm-none-eabi-ld and
the aarch64 equivalents with stubs that fail loudly, then running the whole
check: all 711 sources assemble, all eight profiles compile and all seven
example builds link without any GNU tool being invoked. The GNU default path
still produces an identical image. The diagnostics were confirmed by pointing
the archiver at a name that does not exist and checking that the reason appears.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-10 15:10:34 -04:00
Frédéric Desbiens 68eb205768 Made the Arm ports build with LLVM, and added a check that keeps them that way (#593)
The gnu ports are only ever built with GNU tooling, and GNU as accepts several
non-canonical forms that LLVM's assembler rejects. Nothing noticed, because
nothing built them with anything else. This matters beyond clang itself: Arm
Toolchain for Embedded is LLVM based and is the successor to Arm Compiler 6, so
these are the code paths ac6 users move onto.

Seven files needed changing, none of which alters the emitted code:

LDREX and STREX take no offset in A32 state; the #imm form is Thumb-2 only. GNU
as drops the redundant zero, LLVM rejects it. Removed from the Cortex-A5, A7 and
A9 SMP protect routines.

ARMv8-M Baseline has no flag-preserving MOV immediate, so GNU as already emits
MOVS. Writing MOVS in the two Cortex-M23 sources says what the assembler was
doing anyway. One of them sits in a branch only compiled for the single mode
secure configurations, which is why it had never surfaced.

The Cortex-M0 schedule routine wrote LDR r0, =#0x10000000 with a stray hash,
which its own sibling file already wrote correctly.

The Cortex-M0 system return routine selected the numbered subsection .text 32,
which makes LLVM place the constant pool beyond the range a Thumb-1 PC relative
load can reach. Plain .text fixes it and GNU accepts either form. The reason is
recorded in the file, since 32 files pair a numbered subsection with a literal
pool load and the rest only escape because Thumb-2 and A32 have far more range.

Add scripts/check_clang.sh, which assembles every Arm gnu port source and
compiles the common C sources for one core per architecture profile, and a
clang_check workflow that installs Arm Toolchain for Embedded and runs it. The
toolchain version is pinned and checksum verified, for the same reason the
runner image is pinned.

The port directory to target mapping in that script is explicit rather than
prefix matched. Prefix matching is what makes cortex_a5 also match cortex_a53
and cortex_a55, which are AArch64, and assembling those as ARM32 produces
hundreds of misleading errors; that mistake cost real time while measuring this,
so the reason is recorded next to the table.

Verified with Arm Toolchain for Embedded 22.1.0: 711 of 711 assembly sources
assemble and 185 of 185 C sources compile for all eight profiles, against 6
assembly failures before the change. arm-none-eabi-gcc still assembles all 315
ARM32 sources, so nothing regressed for GNU. The check was confirmed to fail
when any one of the fixes is reverted.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-10 11:40:54 -04:00
Frédéric Desbiens f23a1807a2 Added bash A profile update scripts and restored ports_arch as their source (#592)
The ARMv7-A and ARMv8-A ports are generated by update.ps1, which needs
PowerShell, so the cortex-a job in ports_arch_check ran on a Windows image and
nobody could reproduce it locally on Linux. Add update.sh beside each
update.ps1, with the same cores, compilers, copy sets and patches, and move the
job to the same Linux image as everything else.

The bash scripts were checked against the PowerShell ones by comparing what
each reports as drifted. They agree exactly on the 63 files the Windows job
last reported, and differ on 12 more, which turn out to be a defect in
update.ps1 rather than in the port. Its two .cproject patterns are written as
'value=`"cortex-a7`"' with backticks that survive into the pattern, so that
replacement has never matched, while the neighbouring Cortex-A7.NoFPU pattern
has no backticks and always worked. The result is that the AC6 example builds
for the A5, A8, A9, A12, A15 and A17 cores name cortex-a7 as their CPU while
their FPU string is correct. The bash scripts do what the PowerShell ones
intended, so regenerating corrects those twelve files.

Restore ports_arch as the source for the rest. The implementation of
_tx_thread_smp_time_get from #555 was applied to the twenty four generated SMP
ports and never to ports_arch, which still held MOV x0, #0 with a FIXME
comment, so regenerating would have replaced a working generic timer read with
a stub. That implementation now lives in the source. The remaining differences
are cosmetic and resolve in favour of the source: a trailing blank line in 38
copies of tx_thread_schedule.S and comment spacing in one tx_port.h.

Note that the Cortex-A VFP fix is already present in ports_arch and was never
at risk, contrary to what the description of the port consistency checks change
said before this was measured.

Extend scripts/check_ports.sh to run the A profile generators too, and make it
fail when a generator fails or is missing rather than reporting a clean tree,
which would have been a false pass.

Pin every workflow to ubuntu-24.04. ubuntu-latest already resolves to that
image, so nothing changes today, but a future migration becomes a deliberate
commit rather than something that happens underneath the -m32 builds.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-09 10:55:29 -04:00
Frédéric Desbiens eb4ec4e3b5 Restored ports_arch as the source of truth for the Cortex-M ports (#590)
The Cortex-M ports under ports/ are generated. scripts/copy_armv7_m.sh copies
one tx_port.h and the per tool sources to fifteen M3, M4 and M7 targets, and
scripts/copy_armv8_m.sh does the same for nine M33, M55 and M85 targets. The
ports_arch_check workflow runs both scripts and fails if the tree is not
reproducible, so those copies are meant never to be edited directly.

They were. Every Cortex-M fix since #523 was applied to the generated copies
and not to the source, so the source fell behind and the check went red:
running the three scripts on dev changes 35 files. The check triggers only on
pull requests targeting master, which is why nothing caught it while the fixes
were merged into dev.

Left alone, the next run of these scripts would have reverted three separate
pieces of work: the memory barriers and clobbers from #523, the correction of
the IAR assembly header to use the assembler's own comment syntax, and the move
of tx_initialize_low_level.S into example_build for the M33, M55 and M85 GNU
ports from #514.

Bring the sources up to what the ports carry today, and regenerate. Two
behavioural changes come with that, both deliberate. The barriers from #523
reach the ac5 and keil variants of M3, M4 and M7, which were outside the scope
of that fix and never received it. The barrier that follows restoring the
interrupt posture, which #523 gave only to the GNU ports because GNU was the
only toolchain that could be tested, now applies to every tool; the identical
asm statement already shipped in the AC6 and IAR ports, so this adds a pipeline
flush rather than any new compiler exposure.

Regenerating also drops a stray #endif at the end of the Cortex-M85 IAR
tx_port.h, added by #523, which left that header with one more #endif than #if
and unable to compile. Every other ARMv8-M port was balanced.

Verified that the scripts are idempotent afterwards, that ports_arch_check
would pass, that no port loses a barrier or a clobber, that every regenerated
header is preprocessor balanced, and that every Cortex-M port covered by the
two scripts now carries the entry barrier.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-09 10:42:54 -04:00
Frédéric Desbiens cbdf924d6b Removed the duplicated function body in the Cortex-M4 AC6 port (#589)
_tx_thread_system_return_inline() in the Cortex-M4 AC6 tx_port.h was followed
by a second, orphaned copy of its own body. The copy had no function header, so
it declared interrupt_save at file scope and then placed statements there,
which does not compile. It is also the older version of the body, without the
dsb and isb barriers, so it was left behind rather than intended: the barriers
were added by commit 33efad3f and the previous text was not removed.

Delete the orphaned copy. What remains is the same body every sibling port
carries: after this change ports/cortex_m4/ac6/inc/tx_port.h differs from
ports/cortex_m4/iar/inc/tx_port.h and ports/cortex_m7/ac6/inc/tx_port.h only in
the port name in the banner and the version string, as it should.

Verified by compiling the function in isolation, which fails on the file scope
statements before the change and is clean afterwards, and by a structural scan
of all 208 tx_port.h files in the repository confirming this was the only
occurrence.

Fixes https://github.com/eclipse-threadx/threadx/issues/569

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-09 10:12:09 -04:00
Frédéric Desbiens f84dd3d2aa Added Cortex-R52 port with Armv8-R AEM FVP example build (#579)
Added a ThreadX port for Arm Cortex-R52 with the GNU toolchain

Adds ports/cortex_r52/gnu together with a CMake build, an example build for the
freely available Armv8-R AEM FVP, and an automated test suite.

The port is EL2-aware by construction.  Cortex-R52 always implements EL2 and
resets into it, so the reset path configures EL2, installs both the EL2 and EL1
vector tables, and only then drops to EL1 to run the kernel.  Keeping that
structure from the start lets partitioning work reuse the boot path rather than
replace it; TX_R52_BOOT_AT_EL1 skips the EL2 stage where a vendor monitor has
already dropped privilege.

This is also the first R-profile port in the tree with a working CMake build.
cmake/cortex_a9.cmake exists but ports/cortex_a9/gnu/CMakeLists.txt is empty, so
no A- or R-profile port could be built this way before.

Contents:

  - Kernel port: 16 assembly sources plus tx_port.h, seeded from the
    Cortex-R5/GNU port.
  - Build: cmake/cortex_r52.cmake and the port CMakeLists.txt, with soft/hard
    float, VFP, FIQ and IRQ/FIQ nesting options.
  - Example BSP: EL2 to EL1 boot, GICv3, generic timer, PMSAv8-R MPU,
    semihosting and PL011 consoles.
  - Tests: six images registered with CTest, each judging itself and
    terminating the model.
  - tx_port_offset_check.c: compile-time assertions on the TX_THREAD offsets
    the assembly reaches by hard-coded displacement.  This is the same defect
    class fixed for Cortex-R4/R5 in #578; nothing in the toolchain ties those
    literals to the C structure.

Two facts were established by experiment rather than taken from documentation,
and are recorded in the code and the readme because both are traps:

  - PRBAR.AP bit order is reversed relative to th
    AArch64 macro set: the low bit is read-only and the high bit grants EL0
    access.  Programming four disjoint regions, one per encoding, and
    attempting a privileged write to each gave 0b
    and 0b10 allowed.  Using the AArch64 ordering
    as read-only and silently accept writes, because region coverage is still
    enforced: an unmapped address faults while a "read-only" region does not.
  - The generic timer PPI is INTID 30, recorded from ICC_IAR1 after enabling
    the whole PPI range 25-31, rather than assume

Model behaviour that a silicon port must revisit: the FVP leaves CNTFRQ at zero
with the system counter stopped, so the BSP programs both; CNTHCTL.PL1PCTEN,
CNTHCTL.PL1PCEN and ICC_HSRE must be set at EL2 or EL1 accesses trap; and
tx_thread_vfp_enable() sets only a per-thread software flag, so enabling the FPU
hardware (CPACR, FPEXC.EN) is the board support package's responsibility, not
the kernel's.  The FVP reports MIDR 0x410FD0F0, an architecture envelope model
rather than a Cortex-R52, so it validates architecture and not implementation.
Nothing in the port is gated on MIDR.

Changes outside ports/cortex_r52, three and all small: cmake/cortex_r52.cmake is
a new toolchain file; CMakeLists.txt gains enable_testing() at the top level,
without which add_test() in a subdirectory generat
the root never references, so ctest reports no tests from the build root; and
.gitignore covers build directories and __pycache__.

Verification: all six images build warning-free a
FVP_BaseR_AEMv8R in three configurations -- protection off, protection and
caches on, and both together with hard-float VFP.  Six build-option
combinations build clean and pass the tick-and-preemption demo at run time.

The tests assert behaviour rather than configuration, which is what caught the
problems above: the cooperative demo asserts the order of execution, because
counting alone cannot distinguish working context switches from one thread
running to completion, and the MPU test provokes real faults, because reading
SCTLR back only proves a bit was set.

The 96-test regression suite is host-side and pas
configurations on the Linux port.  It is not cross-run here: it validates
portable kernel logic, while the port-specific risk is the assembly, which the
FVP images exercise.  Structural coverage is therefore not claimed.

Deliberate deviations: passing a string literal to tx_thread_create reports a
discarded const qualifier from common/inc/tx_api.
CHAR *; demo_threadx.c keeps the shipped sample's int main() so the demo body
stays byte-identical to it; and the floating-point test compares exactly on
purpose, since a tolerance would mask a restored
wrong.

Not included: modules support, split-mode SMP and
support. The example targets the FVP only, and NXP S32Z280 bring-up is
separate work; the readme flags what to re-verify there.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-08 08:00:01 -04:00
Frédéric Desbiens 2baae7c57f Fixed VFP issues on Cortex-R4 and R5 ports (#578)
* Fixed missing VFP thread extension on Cortex-R4 and R5 ports

Building cortex_r4/gnu, cortex_r5/gnu or cortex_r5/ac6 with
TX_ENABLE_VFP_SUPPORT corrupts memory. The assembly in
tx_thread_schedule, tx_thread_system_return and tx_thread_context_restore
reads and writes a per-thread VFP enable flag as [thread, #144], but every
TX_THREAD_EXTENSION in these ports' tx_port.h was empty, so no such member
existed.

Offset 144 in the resulting TX_THREAD is tx_thread_filex_ptr, so the
damage runs both ways:

  - tx_thread_vfp_enable() stores 1 into tx_thread_filex_ptr, after which
    FileX dereferences 0x1;
  - conversely, once FileX sets that pointer the context switch reads it as
    "floating point enabled" and starts pushing about 132 bytes of
    floating-point context onto a thread stack never sized for it, then
    restores unrelated data into the D registers.

sizeof(TX_THREAD) is 180 on these ports, so 144 is a live member rather
than padding past the end of the structure.

What makes this dangerous is that it is silent. The defect was reproduced
at runtime by reverting the Cortex-R52 port -- which inherited the same
assembly and the same hard-coded 144 from the R5 port -- to this state and
running its floating-point demo on the Armv8-R AEM FVP. Every
floating-point value was still preserved exactly and no corruption was
reported across interrupts, because a non-zero filex_ptr reads as "enabled"
and the context switch dutifully saves and restores. The only visible
damage was tx_thread_filex_ptr reading 0x00000001. In other words the
feature you enabled appears to work, and what breaks is an unrelated
pointer that nothing notices until FileX is introduced. With the field
present the same demo reports the sentinel intact.

This appears to be drift rather than a deliberate limitation. All fourteen
Armv7-A port and toolchain combinations define the field and contain the
VFP code paths. The R-profile family is inconsistent in both directions:
these three have the code paths without the field, while cortex_r4/ac5,
cortex_r4/ac6, cortex_r4/iar, cortex_r5/ghs, cortex_r5/iar, cortex_r7/ghs
and cortex_r8_smp/ac5 define the field but contain no VFP code paths.

The fix follows the Armv7-A ports exactly: TX_THREAD_EXTENSION_2 carries
tx_thread_vfp_enable, and tx_thread_vfp_enable/disable are declared. The
field is defined unconditionally rather than under TX_ENABLE_VFP_SUPPORT
because the offset is hard-coded in assembly, so a conditional member
would shift every following field and be correct in only one
configuration. No assembly changes are needed: 144 is already correct once
the member exists.

Verified with GCC 14.3: on all three ports tx_thread_vfp_enable now lands
at exactly offset 144, matching the assembly, with tx_thread_filex_ptr
moved to 148. The library builds cleanly for cortex_r4/gnu and
cortex_r5/gnu both with and without TX_ENABLE_VFP_SUPPORT, and the VFP
save and restore instructions appear in the port assembly only when it is
requested. cortex_r5/ac6 is verified by offset and inspection only, as
that toolchain was not available here. Runtime verification was performed
on Armv8-R as described above rather than on R4/R5 silicon; upstream QEMU's
xlnx-zcu102 machine holds its Cortex-R5F cores in reset, so no R5 target
was available.

ABI note for release documentation: adding the member moves
tx_thread_filex_ptr and everything after it by four bytes and grows
TX_THREAD from 180 to 184 bytes. This is the layout the Armv7-A ports
already have, so the change aligns R-profile with A-profile rather than
introducing a new one, but kernel awareness and TraceX tooling that
hard-codes offsets will need rebuilding.

Follow-up worth considering separately: nothing in the toolchain ties the
literal 144 in the assembly to the C structure, which is what allowed this
to go unnoticed. A compile-time assertion per port turns any recurrence
into a build failure -- when the experiment above reverted the field, that
assertion failed the build before a corrupting binary could be produced.
ports/cortex_r52/gnu/src/tx_port_offset_check.c is a working reference.


* Added compile-time offset checks to the Cortex-R4 and R5 ports

Nothing in the toolchain connects the structure offsets hard-coded in the
port assembly to the C definition of TX_THREAD, which is what allowed the
missing VFP thread extension to go unnoticed. These files assert the
offsets that the context-switch path depends on, so a recurrence becomes a
build failure instead of memory corruption discovered later at runtime.

Three offsets are asserted: the VFP enable flag at 144, the thread stack
pointer at 8 and the run counter at 4. The stack pointer and run counter
precede every extension macro in TX_THREAD, so they are stable by
construction. Offset 144 was checked against the build options that could
plausibly move it -- stack checking, event trace, event logging,
performance info and TX_NOT_INTERRUPTABLE -- and is unchanged by all of
them, because the only conditional member nearby guards
tx_thread_filex_ptr, which follows the extension, and
TX_THREAD_USER_EXTENSION appears much later in the structure. The
assertions therefore cannot misfire on a legitimate configuration.

C99 has no _Static_assert, so a negative array dimension is used. Each
file compiles clean under -std=c99 -pedantic -Wall -Wextra and emits zero
bytes of code or data.

Verified that the checks actually catch regressions rather than merely
compiling: asserting a deliberately wrong offset is rejected on all three
ports, and removing the VFP field again fails the build outright with a
message naming the missing member.

These ports have no CMake build of their own, so the files take effect
only in builds that compile everything under src/. That is still
worthwhile given they cost nothing and emit nothing. Extending the same
technique to the remaining ports is deliberately left as separate work: it
requires reading each port's assembly to attribute every hard-coded
literal to the right structure, and copying assertions without that
analysis would risk asserting wrong offsets.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-06 10:05:33 -04:00
Frédéric DesbiensandCodex a33f6efc48 Updated Cortex-R4 Thumb bit immediates (#565)
Use explicit immediate syntax when testing the Thumb bit in AC6 Cortex-R4 stack build assembly.

Co-authored-by: Codex <codex@openai.com>
2026-06-29 16:22:52 -04:00
Frédéric DesbiensandCodex 3fcfb4e4c9 Updated Cortex-M BASEPRI zero immediates (#564)
Use explicit immediate syntax when clearing BASEPRI in GNU and AC6 Cortex-M scheduler assembly.

Co-authored-by: Codex <codex@openai.com>
2026-06-29 16:07:12 -04:00
d9e7dfea89 Add SysTick counter reset in Cortex-M ports (#561)
* Add SysTick counter reset in Cortex-M ports

The SysTick counter value (SYST_CVR @0xE000E018) is indeterminate after
reset. Without clearing it prior to enabling the counter, the first tick
interval becomes unpredictable.

* Fixed missing comment and added generated-by header in Cortex-M0/AC6 port

Added the missing inline comment on the LDR r1, =0 instruction
(Build value for SysTick reset) and the required AI-generated
attribution comment under the copyright header.

---------

Co-authored-by: Frédéric Desbiens <frederic.desbiens@eclipse-foundation.org>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-24 12:16:51 -04:00
franco 859c098747 Fix nested comment warning in RX GCC ports (#550)
The RX GCC assembly port contains an unclosed C-style comment in tx_thread_context_save.S. Because .S files are passed through the C preprocessor before assembly, the nested comment can trigger GCC warnings.

Close the affected comment block without changing assembly instructions or runtime behavior.

Signed-off-by: FranCDoc <fchiesadoc@gmail.com>
2026-06-16 10:26:26 -04:00
Frédéric DesbiensandCopilot df30b8b96e Release 6.5.1.202602a preparation (#547)
* Updated version number constants

* Updated port version strings

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-08 16:59:04 +02:00
Frédéric Desbiens 5a07bdac35 Fix/issue 545 basepri endif (#546)
* Fixed #else replaced with #endif in tx_port.h for cortex_m3/m4/m7 gnu/ac6/iar ports

In the __set_basepri_value/#ifdef TX_PORT_USE_BASEPRI block, an #else
directive was incorrectly replaced with #endif in a prior commit.
This caused the TX_PORT_USE_BASEPRI guard to close prematurely and
left __enable_interrupts unconditionally visible (or triggered an
orphaned #endif / missing #endif error at compile time).

Affected ports: cortex_m3/ac6, cortex_m3/iar, cortex_m4/ac6,
cortex_m4/gnu, cortex_m4/iar, cortex_m7/ac6, cortex_m7/gnu,
cortex_m7/iar.

For cortex_m3/gnu and cortex_m4/gnu the fix restores the #else and
the existing second #endif now correctly closes the block.
For the ac6/iar variants of m3/m4 and all m7 variants, the #else is
restored and a missing closing #endif is added after __enable_interrupts.

Also adds Linux shell build scripts (build_threadx.sh and
build_threadx_sample.sh) for the cortex_m3/gnu, cortex_m4/gnu, and
cortex_m7/gnu example_build directories as Linux equivalents of the
existing .bat files.


* Fixed missing linker script symbols in cortex_m4/gnu example_build

cortexm4_crt0.S references several symbols that were absent from
sample_threadx.ld, causing undefined-reference linker errors when
building the sample:

- __text_load_start__, __text_start__, __text_end__
- __rodata_load_start__, __rodata_start__, __rodata_end__
- __fast_load_start__, __fast_start__, __fast_end__
- __ctors_load_start__, __dtors_load_start__

Changes:
- Added __text_start__/__text_end__ bounds to the .text section and
  __text_load_start__ via LOADADDR(.text).
- Moved .rodata out of .text into its own section with start/end/load
  
- Added __ctors_load_start__ and __dtors_load_start__ markers within
  .text (load address == VMA since the section is XIP in FLASH; the
  crt0 copy call becomes a no-op).
- Added an empty .fast section in RAM with __fast_load_start__ pointing
  to __fast_start__ so the crt0 fast-copy call is a no-op.


* Added Linux build scripts and fixed path bug in cortex_m23 for all ports changed between v6.5.0 and v6.5.1

Added build_threadx.sh (and build_threadx_sample.sh where applicable)
as Linux equivalents of the Windows .bat scripts for all GNU-toolchain
ports touched between v6.5.0.202601_rel and v6.5.1.202602_rel:

  arm9, arm11, cortex_a5/a7/a8/a9/a12/a15/a17,
  cortex_m0, cortex_m23, cortex_r4, cortex_r5

For cortex_m33/m55/m85 (which had no .bat equivalent), build_threadx.sh
was written from scratch using the port CMakeLists.txt source lists and
the correct CPU flags (-mcpu=cortex-m33/m55/m85 -mthumb).

Also fixed a pre-existing bug in cortex_m23/gnu/example_build/
build_threadx.bat (and the generated .sh): tx_thread_stack_error_handler.c
and tx_thread_stack_error_notify.c were referenced as ../src/ (port
directory) instead of ../../../../common/src/ where they actually live.

All 16 newly added build_threadx.sh scripts were verified to compile
successfully with arm-none-eabi-gcc 13.2.1.

---------
Closes #545
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-08 16:51:08 +02:00
Frédéric DesbiensandCopilot 730b61874b Added copyright headers to files missing them
Applied the standard MIT license header to all project-owned C, header,
assembly, shell, and Python files that were missing a copyright notice.
Third-party, toolchain startup, and auto-generated files were excluded.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-06 21:48:06 +02:00
Frédéric DesbiensandCopilot de1c6e9bbe Release 6.5.1.202602 preparation (#543)
* Updated version number constants
* Updated port version strings

---------

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-06-03 16:55:01 -04:00
Frédéric DesbiensandCopilot d4b9448f84 Added riscv-none-elf-rv32imc toolchain file for xPack (risc-v32) (#539)
Add cmake/riscv-none-elf-rv32imc.cmake targeting the xPack
riscv-none-elf-gcc toolchain with rv32imc_zicsr/ilp32 ABI for
bare-metal CORE-V MCU builds.

Update cmake/riscv32-unknown-elf-rv32imc.cmake to correctly target
rv32gc/ilp32d (matching riscv-collab riscv32-elf ABI) for QEMU
regression tests.

Update cmake/riscv64-gcc-rv32imc.cmake compat alias to include
the new riscv-none-elf-rv32imc.cmake.

Update CORE-V MCU example_build: build.sh exports xPack PATH,
install_deps.sh downloads xPack 15.2.0-1, README.md reflects new
toolchain name/path.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-05-27 17:06:37 -04:00
Frédéric DesbiensandCopilot 6061c43f96 Removed clz.c workaround; use riscv32-unknown-elf toolchain (#538)
bsp/clz.c provided a weak __clzsi2 fallback to work around the missing
rv32 multilib in the riscv-collab riscv64-unknown-elf toolchain.  Since
cmake/riscv32-unknown-elf-rv32imc.cmake now uses the dedicated riscv32-
unknown-elf-gcc toolchain (riscv-collab riscv32-elf release), which ships
a native rv32/ilp32 libgcc with all required helpers, the workaround is
no longer needed.

Remove bsp/clz.c and its entry in CMakeLists.txt.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-05-27 12:00:47 -04:00
Frédéric DesbiensandCopilot 80cbaadfaf cmake: rename rv32 toolchain file to riscv32-unknown-elf-rv32imc.cmake (#537)
Renamed cmake/riscv64-gcc-rv32imc.cmake to the accurate name
cmake/riscv32-unknown-elf-rv32imc.cmake and switch the compiler
from riscv64-unknown-elf-gcc to riscv32-unknown-elf-gcc.

The riscv-collab riscv64-elf toolchain has no rv32 multilib and
will fail to link soft-float and integer helpers (__clzsi2, __muldf3,
etc.) when building for -march=rv32imc_zicsr -mabi=ilp32.  The
dedicated riscv32-unknown-elf-gcc (riscv-collab riscv32-elf release,
installed to /opt/riscv by scripts/install_riscv.sh) ships the correct
native rv32/ilp32 libgcc — analogous to arm-none-eabi-gcc for Cortex-M.

The old filename is kept as a two-line compatibility alias that includes
the new file, so any out-of-tree users who hardcode the old path still work.

Also update:
- core_v_mcu/build.sh: reference new cmake filename
- core_v_mcu/README.md: update prerequisites table and toolchain docs

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-05-27 11:33:28 -04:00
Frédéric DesbiensandCopilot 223556219+Copilot@users.noreply.github.com 7486de06c8 Refactored, consolidated, and cleaned up RV32/RV64 ports (#536)
risc-v: refactor, consolidate, and fix RV32/RV64 ports

Consolidates the RISC-V 32-bit and 64-bit GNU/Clang port sources, fixes two
pre-existing assembly bugs discovered during testing, and hardens the build
infrastructure for both the regression suite and the CORE-V MCU example.

--- Port consolidation (RV32 GNU + Clang) ---

 - Delete ports/risc-v32/clang/src/ (8 .S files had no Clang-specific
 directives; diverged from GNU only due to missing bug fixes). The Clang
 port CMakeLists.txt now compiles from ../gnu/src/.
 - Change .global -> .weak for _tx_initialize_low_level in gnu/src/ to allow
 BSP-level override without a linker conflict (adopted from Clang port).
 - Create ports/risc-v32/common/tx_port_riscv32_common.h with all definitions
 shared between GNU and Clang ports. Reduce both tx_port.h files to thin
 wrappers.
 - Add a prominent comment in risc-v64/gnu/inc/tx_port.h explaining why
 LONG/ULONG are intentionally 32-bit on RV64 (ThreadX ABI requirement,
 mirrors win64/MSVC LLP64).

--- Shared CMake helper ---

 - Add cmake/threadx_riscv_port.cmake with threadx_add_riscv_port(). All
 three port CMakeLists.txt files are reduced to ~8 lines each. Include path
 is relative to CMAKE_CURRENT_LIST_DIR so the helper works whether ports
 are built standalone or as a subdirectory of the test framework.

--- Shared example-build drivers ---

 - Create canonical driver files under ports/risc-v_common/:
  inc/csr.h                  (uintptr_t-based; portable RV32 + RV64)
  example_build/plic/        (plic.c, plic.h)
  example_build/uart/        (uart_qemu_ns16550.c/h; static inline putc_nolock)
  example_build/trap/        (trap_qemu.c; XLEN-portable mcause constants)
 - Replace per-example copies with symlinks in all qemu_virt and cva6_ariane
 example directories.
 - Fix OS_IS_INTERRUPT typo (was OS_IS_INTERUPT) in shared trap_qemu.c.
 - Gate print_hex() behind TX_RISCV_TRAP_DEBUG.

--- Bug fixes in RV32 assembly ---

tx_thread_schedule.S:

 - Solicited-return FP path: reload t0 from the mepc stack slot before
 csrw mepc, t0. After the FP restore block, t0 held the fcsr value (0 for
 new threads), which caused mepc = 0 and an immediate instruction-address
 fault on the first context switch.
 - Same path: reload t0 from the mstatus stack slot before csrw mstatus, t0
 to avoid writing the stale fcsr value into mstatus.

tx_thread_system_return.S:

 - FP callee-saved registers were saved unconditionally before the mstatus.FS
 check, causing an illegal instruction trap (mcause=0x2) when a thread with
 FS=Off (lazy FPU, thread has never used FP) voluntarily yielded.
 - Apply the same FS guard pattern used in tx_thread_context_save.S: read
 mstatus first, isolate FS[1:0], and skip fsw/fsd if FS == Off.

Both bugs were pre-existing on origin/dev and are unrelated to the
consolidation changes.

--- RV64 64-bit pointer compatibility ---

 - Add TX_TIMER_INTERNAL_EXTENSION, TX_THREAD_CREATE_TIMEOUT_SETUP, and
 TX_THREAD_TIMEOUT_POINTER_SETUP to risc-v64/gnu/inc/tx_port.h to store the
 thread timeout pointer in a VOID
  * extension field rather than truncating it
 into a 32-bit ULONG. Mirrors the win64 port pattern.
 - Define TX_TIMER_EXTENSION_PTR_DEFINED as a portable sentinel.
 - Update threadx_thread_basic_execution_test.c guard from #if defined(_WIN64)
 to #if defined(_WIN64) || defined(TX_TIMER_EXTENSION_PTR_DEFINED).
 - Disable -Wconversion for the RV64 test build: ULONG = unsigned int (32-bit)
 is intentional for ThreadX ABI but triggers spurious warnings when sizeof()
 (8 bytes on RV64) appears in arithmetic with ULONG in common/src/.

--- Regression suite cmake fixes ---

test/tx/cmake/riscv/regression/CMakeLists.txt:

 - Build testcontrol_weak_defaults.c as a separate OBJECT library and include
 it in every test executable via $<TARGET_OBJECTS:>. GNU ld does not extract
 objects from a static archive to satisfy weak symbols, so bundling it in
 test_utility was insufficient for the standalone
 threadx_initialize_kernel_setup_test.

test/tx/cmake/regression/CMakeLists.txt,
test/smp/cmake/regression/CMakeLists.txt:

 - Same fix applied to the Linux and SMP regression builds. The symbols
 abort_all_threads_suspended_on_mutex, suspend_lowest_priority, and
 abort_and_resume_byte_allocating_thread were introduced by the win64 merge
 and left the standalone test unlinkable.

--- CORE-V MCU toolchain and build fixes ---

cmake/riscv64-gcc-rv32imc.cmake:

 - Resolve riscv64-unknown-elf-gcc via PATH so the riscv-collab toolchain in
 /opt/riscv/bin is preferred when it appears first.

ports/risc-v32/gnu/example_build/core_v_mcu/bsp/clz.c (new):

 - The riscv-collab toolchain is built without rv32 multilib, so its libgcc
 does not define __clzsi2 (the helper emitted for __builtin_clz() in fll.c).
 Add a weak __clzsi2 fallback so the build is self-contained with any
 riscv64-unknown-elf toolchain. The weak attribute yields to a
 libgcc-provided strong symbol when the Ubuntu multilib package is used.

core_v_mcu/CMakeLists.txt:

 - Add bsp/clz.c to sources.
 - Reference CMAKE_TOOLCHAIN_FILE via message(STATUS) to suppress the false-
 positive "Manually-specified variables were not used by the project" CMake
 warning and to show the active toolchain at configure time.

--- Housekeeping ---

 - Rename azrtos_test_* -> threadx_test_* (eliminate Azure RTOS branding).
 - Add RV64 QEMU CI test script:
  ports/risc-v64/gnu/example_build/qemu_virt/test/
  threadx_test_tx_gnu_riscv64_qemu.py
 - Normalize entry.s -> entry.S in all 4 example directories.
 - .gitignore: exclude build_m7/ and .codex local artifacts.
 - CI: comment out the riscv regression workflow job and remove it from the
 deploy job's needs list (preserved in-place for easy re-enablement).

--- Verified ---

 - 95/95 RV32 regression tests pass (QEMU virt)
 - 95/95 RV64 regression tests pass (QEMU virt)
 - All 5 Linux build configurations build cleanly (default_build_coverage,
 disable_notify_callbacks_build, stack_checking_build,
 stack_checking_rand_fill_build, trace_build)
 - CORE-V MCU example_build links cleanly with /opt/riscv toolchain

Co-authored-by: Copilot 223556219+Copilot@users.noreply.github.com
2026-05-27 10:30:57 -04:00
Frédéric Desbiens 2c16114a45 Added win64 ports of ThreadX and ThreadX SMP (#529)
Windows x64 port and regression suite

 This PR adds the Windows x64 (Win64) simulation port for both the standalone
 and SMP variants of ThreadX, along with the full CMake build and test
 infrastructure needed to run the regression suite on Windows.


 New ports

 Win64 standalone (ports/win64/vs_2022): self-contained Windows simulation
 port using Win32 threading primitives as virtual cores. Includes CMake
 integration, build/test scripts, and MSVC project files.

 Win64 SMP (ports/win64_smp/vs_2022): multi-core Windows simulation port.
 Supports up to 4 virtual cores backed by Windows host threads.


 Scheduler and timer improvements

 The initial port used coarse polling and synchronous SuspendThread/ResumeThread
 pairs throughout the scheduler hot path. Several rounds of optimization reduced
 the SMP regression suite runtime from ~150 s to ~78 s (-48%), with no
 regressions:

 - Replaced scheduler polling with an event-driven wake path; switched the
   simulated timer to one-shot rearming to eliminate catch-up ticks.
 - Skip SuspendThread when _tx_thread_preempt_disable != 0 (new suspension
   type 3) -- the primary optimization, yielding up to 7.9x speedup on
   preemption-heavy tests.
 - Skip SuspendThread when a thread is spinning on the Win32 critical section
   (suspension type 4), and fix a stale-TLS bug in
   _tx_win32_critical_section_obtain that could stamp mutex_access on the
   wrong virtual core.
 - Added a 2 ms scheduler event timeout (matching the Linux SMP port) to
   prevent stalls on any missed SetEvent.
 - Enabled high-resolution waitable timers (SetWaitableTimerEx) for accurate
   100 Hz tick cadence.
 - Increased TX_WIN32_CONTENTION_PAUSE_COUNT from 64 to 256 to reduce
   SwitchToThread overhead under heavy CS contention.


 Build and test infrastructure

 - Hardened the Windows build wrapper (scripts/build_tx.ps1): invoke Ninja
   directly for Ninja build trees, fix timeout detection, add a default build
   timeout, and limit fallback replay to real timeout cases.
 - Added -Clean support to Windows test scripts to remove stale CTest state
   before each run.
 - Skip Visual Studio DevShell re-entry when the active MSVC environment
   already matches the requested architecture.
 - Fixed scripts/build_tx.sh (Linux) regression source generation: replaced
   brittle exact-string insertion with line-based matching so the interrupt
   dispatcher hook is inserted reliably for both simulator ports.


 Test suite updates

 - Introduced test/tx/regression/threadx_test_port.h with portable macros
   (TX_TEST_POINTER_WORD, TX_TEST_STORE_POINTER) for storing pointers in test
   arrays on 64-bit targets where ULONG remains 32-bit.
 - Adjusted pool-capacity and pointer-storage patterns in regression tests to
   use ALIGN_TYPE-sized slots, making the suite correct on 64-bit hosts.
 - Restored stricter event flag, sleep, and timer expectations now that
   port-level fixes make prior Windows accommodations unnecessary.
 - Tightened SMP watchdog and clean-build timeout defaults.


 Version metadata

 Updated Win32, Win64, and Win64 SMP port version strings to 6.5.1.202602.

 Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
 Co-authored-by: Codex (gpt 5.5) <codex@openai.com>
2026-05-26 17:17:40 -04:00
Wei-Chen LaiWei-Chen Lai Winstonllllai@users.noreply.github.comFrédéric Desbiens frederic.desbiens@eclipse-foundation.orgCopilot 223556219+Copilot@users.noreply.github.com
e031a3d506 Implemented lazy FPU, GP relaxation, and QEMU automation for GNU port in arch/risc-v32 (#513)
This PR adds three functional improvements to the RISC-V 32-bit GNU port.

Lazy FPU stacking (tx_thread_context_save.S, tx_thread_context_restore.S, tx_thread_schedule.S): FP register save/restore is now skipped whenmstatus.FS is Off, reducing context switch overhead for threads that do not use floating point.

GP relaxation (cmake/riscv32_gnu.cmake, entry.s, link.lds): Enables the -mrelax compiler flag and defines __global_pointer$ in the linker script.The entry stub initializes gp at startup. gp is not saved or restored during context switches.

WFI in idle loop (tx_thread_schedule.S): The scheduler issues wfi when no thread is ready, replacing busy-waiting with a low-power sleep.

A Python/QEMU/GDB functional test runner is added under ports/risc-v32/gnu/example_build/qemu_virt/test/. It validates context switching, FPUcontext preservation, timer interrupts, and preemption. To run:

cd ports/risc-v32/gnu/example_build/qemu_virt
make check-functional-riscv32

Tested on QEMU virt machine (rv32gc).

Co-authored-by: Wei-Chen Lai Winstonllllai@users.noreply.github.com
Co-authored-by: Frédéric Desbiens frederic.desbiens@eclipse-foundation.org
Co-authored-by: Copilot 223556219+Copilot@users.noreply.github.com
2026-05-25 16:43:24 -04:00
Christos Papadopoulos 5b1b740a66 Reverted ULONG to be 4 bytes long, fixing missaligment issue in the current port (#534)
* Reverted RISCV-64 port

* Fixed wrong load/store instructions used
2026-05-25 16:12:33 -04:00
c7d4b2f228 Added RISC-V board Bananapi BPI-F3 (SpacemiT K1 SoC) BSP Support (#531)
* Add BananaPi BPI-F3 BSP support

Add RISC-V supervisor support to the rv64/gnu port and
provide a complete board support package for the BananaPi BPI-F3
(SpacemiT K1 SoC, X60 cores).

Port changes (risc-v64/gnu):
- Guard all CSR accesses with TX_RISCV_SMODE to select S-mode registers
  (sstatus/sepc/sie/sret) vs M-mode (mstatus/mepc/mie/mret) in
  context_save, context_restore, schedule, system_return,
  interrupt_control, and stack_build.
- Add S-mode TX_INT_ENABLE/TX_DISABLE and inline TX_RESTORE macros
  to tx_port.h.
- Add TX_RISCV_SMODE CMake option to CMakeLists.txt.

BananaPi BPI-F3 BSP (example_build/bananapi-f3):
- Boot flow: FSBL → OpenSBI (M-mode) → U-Boot (S-mode) → ThreadX
- S-mode trap handler with context save/restore integration
- SBI legacy ecall timer at 10 Hz (24 MHz timebase)
- PLIC driver with S-mode context, stale-IRQ drain, and callbacks
- PXA-compatible UART0 console (115200 8N1)
- Linker script at 0x200000 load address

Tested on risc-v board, BananaPi BPI-F3 hardware.

Signed-off-by: Akif Ejaz <akif.ejaz@10xengineers.ai>

* Fixed S-mode context restore and PLIC spurious IRQ handling

- tx_thread_context_restore.S (non-nested path): set SPIE(0x20) alongside
  SPP(0x100) so sret re-enables interrupts in the restored thread.
- tx_thread_context_restore.S (both S-mode paths): use FS=Dirty (0x6000)
  instead of FS=Initial (0x2000) to match tx_thread_schedule.S and prevent
  FP register corruption across context switches.
- plic.c (plic_irq_intr): return early when plic_claim() yields 0 (no
  pending interrupt) to avoid completing a spurious IRQ ID 0, which is
  undefined behavior per the PLIC spec.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

---------

Signed-off-by: Akif Ejaz <akif.ejaz@10xengineers.ai>
Co-authored-by: Frédéric Desbiens <frederic.desbiens@eclipse-foundation.org>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-05-25 15:59:17 -04:00
Frédéric Desbiens 5513b1724a Added port for the OpenHW CORE-V MCU platform (#535)
Adds a complete ThreadX port for the OpenHW CORE-V MCU SoC, targeting the Digilent Nexys A7 FPGA board with an Ashling Opella-LD debug probe.

 New files
 ---------
 cmake/riscv64-gcc-rv32imc.cmake
   CMake toolchain file for riscv64-unknown-elf-gcc targeting rv32imc_zicsr/ilp32.

 ports/risc-v32/gnu/example_build/core_v_mcu/ -- Full BSP + demo application:

Assembly
     crt0.S                     C runtime startup (BSS clear, GP/SP init, call main)
     vectors.S                  32-entry vectored interrupt table at 0x1c000800
     tx_initialize_low_level.S  mtvec setup (vectored mode), stack/free-mem pointers

BSP drivers (bsp/)
     system_core_v_mcu.c  Top-level init and ISR dispatcher (isr_table[32])
     irq.c                PULP APB interrupt controller (enable/disable/mask)
     timer_irq.c          PULP FC Timer -- 100 Hz tick from 10 MHz SOC clock
     fll.c                Frequency Locked Loop -- 5 MHz to 50 MHz (FPGA)
     uart_driver.c        UDMA UART channel 0 -- polled TX, non-blocking RX (8-bit transfer width)
     gpio.c               PULP apb_gpiov2 driver: set/clear/toggle/direction by pin index;
                          pad-mux via per-pad indexed registers at APB_SOC_CTRL + 0x400 + pad*4;
                          full pinmux API (gpio_setpinmux, gpio_getpinmux, gpio_pin_set_dir,
                          gpio_pin_read_status)
     i2c_master.c         Polled UDMA I2C master
     adt7420.c            ADT7420 temperature sensor driver
     string.c             Freestanding memset/memcpy shim (no newlib)

Headers (include/)
     Peripheral register maps, MMIO inlines, BSP API declarations, tx_user.h

Application
     demo_threadx.c   Two threads: LED[0] blink at 1 Hz (IO pad 11, GPIO pin 4, MUX=2)
                      + UART heartbeat with startup banner
                      ("Eclipse ThreadX for OpenHW CORE-V MCU vX.Y.Z.BBBBB")
     link.ld          Linker script: .vectors@0x1c000800, .text@0x1c000880
     CMakeLists.txt   Build definition (references THREADX_ROOT)
     build.sh         One-shot CMake+Ninja build script

Tooling
     install_deps.sh               Automates toolchain/OpenOCD dependency setup
     deploy.sh                      One-step GDB flashing via Ashling Opella-LD;
                                           gdb-multiarch fallback when riscv64-unknown-elf-gdb is absent;
                                           supports --wsl flag required by usbipd-win v5.x
     openocd-nexys-Ashling-Opella-LD.cfg   OpenOCD config for Opella-LD over JTAG

Tests (tests/)
     test_irq.c / test_timer.c   Host-compiled unit tests (2/2 pass)
     mock/mmio_mock.*             Software MMIO register map for host testing

Documentation
     README.md   Hardware overview, build, flash/debug, BSP API reference

 Architecture notes
 ------------------
 - CV32E40P uses the PULP/PULPissimo interrupt controller (not CLINT); IRQ lines
   are masked via APB registers, not the mie CSR.
 - mtvec must be 256-byte aligned; vectors placed at 0x1c000800 (vectored mode).
 - Timer IRQ = line 10; dispatch via isr_table[mcause & 0x1f].
 - Build: -march=rv32imc_zicsr -mabi=ilp32, -ffreestanding, -nodefaultlibs.
 - Verified: ELF 11 KB text, sections at correct addresses, unit tests pass.

 Third-party attributions
 ------------------------
 BSP files derived from core-v-freertos (Apache-2.0):
   (c) 2019-2020 ETH Zurich and University of Bologna
   (c) 2020 GreenWaves Technologies
   (c) 2011-2014 Wind River Systems, Inc.
   (c) 2017 SiFive Inc. (crt0.S, BSD-2-Clause portions)
 All original copyright notices retained; see individual file headers.
 SPDX: Apache-2.0 AND MIT (crt0.S: (Apache-2.0 OR BSD-2-Clause) AND MIT).

 Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-05-25 11:28:55 -04:00
Ian Thompson 9ccdd5d822 Updated Xtensa support (#525)
* Add LX8 support for > 32 interrupts

Also fix inconsistent SWPRI define in interrupt handler

* ThreadX: Fix context switch logic

- Disable interrupts prior to allocating a large exception frame;
  only reenable them after deallocating the extra memory.
- Any user tasks with stacks based on TX_MINIMUM_STACK do not
  have sufficient space for both a context switch frame and an ISR
  frame; ill-timed interrupts were causing stack overruns.

* ThreadX: Interrupt fixes for TX_ENABLE_EXECUTION_CHANGE_NOTIFY

- Must reload register trashed by notify hook function call
- Must ensure PS.WOE is set before using call8
- Remove unused XT_USE_INT_WRAPPER define and associated changes,
  which had a bug in XEA2 usage
- Fix another case where enabling the thread notify hooks for
  call0 ABI corrupted a register

* ThreadX: Support for __DYNAMIC_REENT__

- ThreadX change to handle dynamic reent for both newlib and xclib
- Adapt ThreadX xclib interface code to handle dynamic reent pointers

* ThreadX: Xtensa execution profiling support

- Update upstream execution profiling for Xtensa port

* ThreadX: Update xtensa port readme

* ThreadX: Add Xtensa example to EPK

- tx_execution_profile.h now defines an Xtensa example
  in addition to the existing Cortex example

* ThreadX: Add xtensa.cmake

* ThreadX: Update XSHAL_CLIB ifdefs in __getreent()

- Prevent a fall-through with no return value when neither
  xclib nor newlib are used.
2026-05-19 11:11:37 -04:00
Akif Ejaz c1e3678797 Add QEMU based CI regression test infra for RV32 and RV64 (#526)
Added a QEMU virt-machine BSP and CTest infrastructure to run the
ThreadX regression suite on both RISC-V 32-bit and 64-bit targets
in CI.

New components:
- BSP (entry, trap, PLIC, CLINT timer, UART, linker script) targeting
  QEMU virt machine for RV32 and RV64
- CMake build system with Ninja, supporting multiple build configs
- CI scripts: install_riscv.sh (toolchain + QEMU), build_tx_riscv.sh,
  test_tx_riscv.sh
- GitHub Actions workflow job for RISC-V regression gating

Port fixes:
- RV32 tx_thread_context_restore.S: set MPIE alongside MPP (0x1800 →
  0x1880) so mret re-enables interrupts
- RV32/RV64 tx_port.h: add TX_REGRESSION_TEST extension macros needed
  by the test harness
- RV32/RV64 example_build scripts: add compile and QEMU launch steps

Regression test portability fixes:
- Block memory tests: increase pool sizes (320 → 340) to accommodate
  larger RISC-V block-header alignment
- Byte memory test: replace hardcoded offsets with BYTE_POOL_OVERHEAD
  macro for portable pool-size computation
- Event flag timeout test: make counter tolerance unconditional,
  removing linux-only guard

Signed-off-by: Akif Ejaz <akif.ejaz@10xengineers.ai>
2026-05-19 11:03:23 -04:00
Christos Papadopoulos 0dd6d28afb Fixed MPIE being cleared leading to have interrupts being disable after an mret instruction (#522)
* Fixed MPIE being cleared leading to have interrupts being disable when returning from machine mode

* Removed wrong register operand and replaced by immediate value
2026-05-13 09:00:54 -04:00
Francisco Manuel Merino Torres 42adda4457 Added a riscv32 for the OpenHW CVA6 (#511) 2026-05-13 07:46:54 -04:00
Frédéric DesbiensandCopilot 33efad3fee Fixed race condition and message loss in Cortex-M ports (#523)
* Fixed race condition and message loss in Cortex-M GNU, AC6, and IAR ports (#516)

- Added compiler memory barriers to BASEPRI management functions in tx_port.h.
- Added architectural barriers (DSB/ISB) to scheduler return paths in tx_port.h and tx_thread_system_return.S to prevent fall-through before context switch.
- These changes address spurious thread resumption and lost messages, especially when TX_NOT_INTERRUPTABLE is enabled.
- These changes ensure that pending interrupts (specifically PendSV) are recognised before subsequent instructions are executed, following Kairalite's feedback and ARM architectural guidelines.

 Assisted-by: Gemini (Gemini 2.0 Flash)

-----

* Added a comment in common/tx_queue_cleanup to document why the NI path omits revalidation guards

- In `TX_NOT_INTERRUPTABLE` mode, the caller keeps interrupts disabled across the entire cleanup call, so the race window that makes the guards necessary in the interruptable path cannot occur. Add a comment explaining this, and noting that all paths that resume a suspended thread clear tx_thread_suspend_cleanup before calling
_tx_thread_system_ni_resume, making double-cleanup impossible.

This prevents future false-positive suggestions (e.g. from AI tools) to add redundant checks to the NI path.

Relates to: eclipse-threadx/threadx#516

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-04-29 09:10:57 -04:00
Frédéric Desbiens 454ab56fce Merge pull request #278 from Polarisru/master
Fixed SysTick initialization problem in tx_initialize_low_level.s for Cortex-M0.
2026-04-15 11:23:57 -04:00
Frédéric Desbiens 98e4754381 build: add conditional CMake support for ThreadX SMP
- Introduce THREADX_SMP option in root CMakeLists.txt.
- Implement conditional source and port directory selection for SMP builds.
- Add CMake support for common_smp and Cortex-A9 SMP port.
- Fix linker flags in Cortex-A9 SMP sample build script.
- Remove duplicate invalidateCaches_IS declarations in v7.h headers.

Assisted-by: Gemini (Experimental)
2026-04-15 10:52:14 -04:00
Frédéric Desbiens a0dc185f66 Merge pull request #508 from goodnorning/feature/rv64_rvv_support
Added rv64 rvv support
2026-04-14 09:26:33 -04:00
shuta.lst dfafc96149 RISC-V64 qemu_virt example support RVV Extension; 2026-04-07 14:03:24 +08:00
shuta.lst 2f1fc52918 RISC-V64 arch. port support RVV Extension; 2026-04-07 11:41:33 +08:00
Frédéric Desbiens c3259a2160 Updated copyright headers and version number constants (#509)
* Updated version number constants

* Removed revision history from all files

* Added Eclipse ThreadX contributors' copyright header
2026-03-05 10:46:30 +01:00
Frédéric Desbiens 385d39f421 Merge pull request #502 from quintauris-tech/riscv32-clang-port
Added a RV32 Clang port
2026-03-04 08:36:53 +01:00
Frédéric Desbiens b3e8a9aa20 Merge pull request #500 from goodnorning/feature/support_xuantie_e906
Added support for the XuanTie E906 CPU.
2026-03-04 08:34:59 +01:00
Frédéric Desbiens 0c92f48d50 Merge pull request #493 from mehmetteren/dev
Fixed VFP build failure in Cortex-A tx_thread_schedule.S
2026-03-02 11:41:53 -05:00
Frédéric Desbiens da5093f84b Proposed changes to the original PR 2026-03-02 14:24:07 +01:00
Akif Ejaz cd101d9ab4 update comments
Signed-off-by: Akif Ejaz <akif.ejaz@10xengineers.ai>
2026-03-01 15:13:28 +05:00