Measured context-switch cost against the memory holding the stack (#635)

#634 placed a thread stack in BTCM and deliberately claimed no timing benefit,
because none had been measured. This measures it.

Two pairs of equal-priority threads hand control back and forth with
tx_thread_relinquish. One pair has both stacks in BTCM, the other in DRAM0, and
the measuring thread of each pair times the round trip in PMU cycles. Both pairs
run in one image from one copy of the measuring code, which is what makes the
comparison safe: the alignment trap that invalidated earlier work here bites
when two builds with different layouts are compared, and a code shift moves both
pairs equally.

Reproducible to the cycle across runs:

                        min    mean    max
    BTCM stacks        1370    1379   1402
    DRAM0 stacks       1476    1480   1508

BTCM is about 7.4% faster, or 110 cycles on a round trip of two switches.

Three findings that bound the claim, and the last one deflates it.

The figure holds whether the cache is warm or cold. Cleaning and invalidating
the data cache before every timed switch costs both configurations about 40
cycles and leaves the gap at 7.4%: warm it is 1334 against 1440, cold 1370
against 1476. So the advantage comes from BTCM's zero wait states, not from
avoiding cache misses.

That is because a context switch touches almost no stack -- sixteen registers,
about one cache line -- so the stack's cache state has little to contribute
either way. TCM should matter much more for threads with deep call chains or
large locals, where the stack working set is big enough for cache state to
dominate. That is not measured here and should not be assumed.

There is no determinism benefit visible in this test. Excluding preempted
samples, jitter is 32 cycles for BTCM and 28 to 34 for DRAM0 -- comparable, not
better. A first version of this measurement appeared to show BTCM with 16 times
less jitter, and that was wrong: max was reporting whichever pair a timer tick
had landed on. Across three runs the outlier appeared in the BTCM pair once and
the DRAM0 pair twice. Samples past 2000 cycles are now counted separately and
excluded from min, mean and max alike, and the count is printed so the reader
can see how many there were.

Also tried and discarded: loading the partner thread with a cache walk to create
pressure. The timed round trip includes the partner, so the walk dominated every
sample and put all 256 past the outlier threshold. The per-sample flush replaced
it and sits outside the timestamps.

The demo also starts the PMU cycle counter, which bsp_boot.c does for the probe
image and this image never ran. Without it every reading would have been zero,
which reads as a free context switch rather than as a dead counter.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
This commit is contained in:
Frédéric Desbiens
2026-08-17 17:08:29 -04:00
committed by GitHub
parent 62e966f6bf
commit 367d91880b
File diff suppressed because it is too large Load Diff