Files
threadx/ports
Frédéric Desbiens 62e966f6bf Enabled BTCM, measured it, and put a ThreadX thread stack in it (#634)
* Enabled BTCM and measured what it offers as a data store

BTCM was disabled because of a regression that turned out not to exist; the
claim was retracted in the previous commit. It is enabled now, and this
measures why that is worth doing: 16 KB at zero wait states, where ATCM has
one, and no cache in the path at all.

Enabling it needs three things, all of which existed for ATCM already: the
region register write at EL2, an MPU region, and an ECC preload before any read
(TRM 6.2.2). BTCM accepts 32-bit stores where ATCM needs 64-bit, which
tcm_preload already handles.

The cache benchmark now sweeps three memories rather than one, four loop
alignments each. At the alignments where the loop is not instruction-fetch
bound:

    memory              cold (uncached)   warm (cached)   gain
    DRAM2 half-speed            857,540         651,436   24.0%
    DRAM0 full-speed            797,824         651,297   18.3%
    BTCM  zero wait             694,689         651,369    6.2%

Three things follow.

Warm times are identical across all three memories, within 0.02%. Once the data
cache is working the backing store barely matters, because the working set fits
in it.

Cold times rank as the reference manual predicts: BTCM fastest, then DRAM0,
then DRAM2 at half the core frequency (S32Z2 RM 6.3.6).

BTCM still shows a 6.2% gain when the caches are enabled, and that cannot be
the data cache, because an enabled TCM is Non-cacheable Non-shareable Normal
memory whatever the MPU says. It is the instruction cache on the timing loop.
This probe has always measured both caches together; three memories side by
side is what makes that visible.

The number that matters for placing data in BTCM: uncached BTCM is within 6.6%
of the best cached case, where uncached DRAM0 is 22% off it. Data in BTCM runs
at close to cache-hit speed with no cache to miss, which is the determinism
argument stated as a measurement rather than an assertion.

DRAM2's sweep is unchanged with BTCM enabled, 0 and 0 and 240 and 240, which
independently confirms the retraction.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Put a ThreadX thread stack in BTCM

The code side of TCM was done in #630; this is the data side. One of the demo's
three thread stacks now lives in BTCM and the other two stay in DRAM0, so a run
exercises both paths and a mistake in either shows up.

link.lds gains a BTCM region and a .btcm_bss NOLOAD section, so a stack is an
ordinary C array with a section attribute and the linker checks it fits, rather
than a hardcoded address that silently overflows the bank.

entry.S preloads the whole bank at EL2, and that is not optional. ECC is enabled
on this part, so a TCM location must be written before it can be read (TRM
6.2.2), and a stack is read before the program writes it -- the first context
restore pops what tx_thread_create built into it. The preload has to happen
before any C runs, because the demo images do not run bsp_boot.c, which is where
the ATCM preload lives. 32-bit stores suffice for BTCM where ATCM needs 64-bit.

Why BTCM for a stack: 16 KB at zero wait states where ATCM has one, and never
cached whatever the MPU says about it. Measured in the previous commit, uncached
BTCM comes within 6.6% of the best cached case while uncached DRAM0 is 22% off
it, so stack access runs at close to cache-hit speed without depending on a line
being resident. That is the property a determinism argument needs.

What this commit does not claim: no thread-level timing improvement has been
measured. The case for BTCM here rests on the memory characterisation and on
removing the cache from the path, not on a measured context-switch figure. That
measurement is worth doing and has not been done.

Verified on the S32Z280-594EVB. The demo reports its stack addresses so the
placement is visible rather than implied -- sleeper at 0x30100000 in BTCM,
spinner and judge in DRAM0 -- and passes with 100 ticks, 20 sleeper wakeups and
20 preemptions, so a real thread schedules, preempts and context-switches on a
tightly-coupled-memory stack. The boot image still passes six of six probes.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-17 16:33:55 -04:00
..
…
…
…