Measured the cache benchmark at four alignments, because one is not enough (#631)

This benchmark was bimodal and reported a single number, which made it worse
than no benchmark. The same workload on the same silicon reports either 24%
cache benefit or none at all, decided only by where the loop falls inside a
64-byte line -- and therefore by any unrelated change that shifts code
ahead of it. Two nop instructions added to entry.S were enough to flip it.

That is not a hypothetical. A run of conclusions drawn from this probe turned
out to be measuring code layout: an interrupt-handler comparison, a claim that
enabling a second TCM bank costs all cache benefit, and a follow-up claim that
what mattered was when the TCM region register was written rather than what it
contained. The last of those was reported to NXP as a defect and has had to be
withdrawn. Enabling CTCM and adding two nops produce identical results, to the
digit, because the only measurable consequence of the enable was the eight
bytes of instructions it added.

Four copies of the loop are now generated at different offsets within a cache
line, all four are measured, and the low and high gains are both reported.
Pinning a single alignment was tried first and is not a fix: it silently picks
one of the two modes -- aligned to 64 the loop sits permanently in the low one.

Measured on the S32Z280-594EVB, reproducing exactly across runs:

    loop offset in line    cold      warm      gain
    0                      890,302   890,035   0%
    16                     890,208   889,976   0%
    32                     857,439   651,390   24.0%
    48                     857,631   651,571   24.0%

The cold pass differs between the modes as well, 890k against 857k, so the
loop is slower even with both caches off. The cold pass is instruction-fetch
bound out of code RAM at half the core frequency (S32Z2 RM 6.3.6), and how the
loop straddles lines decides how much of the data cache's contribution is
visible at all. This probe therefore measures both caches together and always
did; the sweep at least makes the variation visible instead of letting one
arbitrary placement stand in for the part.

C4 now passes if any alignment shows a 10% speedup, and says so explicitly
when the low mode does not, so the sensitivity appears in the log rather than
being discovered later.

Verified: with the sweep in place, adding 0, 8, 12 or 20 bytes of nops to
entry.S leaves the reported low and high gains unchanged. Before it, the same
shifts read 24.0%, 0%, 0% and 0%.

Also adds cache_disable_all, which the sweep needs: cache_enable was one-way,
so a second cold reading in one run was impossible.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
This commit is contained in:
Frédéric Desbiens
2026-08-17 13:14:16 -04:00
committed by GitHub
parent b2983c08d3
commit 1d3f3f8e4c
4 changed files with 190 additions and 67 deletions
File diff suppressed because it is too large Load Diff
@@ -318,3 +318,29 @@ unsigned int cache_enabled(void)
return (((sctlr & SCTLR_C) != 0UL) && ((sctlr & SCTLR_I) != 0UL)) ? 1U : 0U;
}
void cache_disable_all(void)
{
unsigned long sctlr;
/* Clean before disabling, so dirty lines reach memory while the cache is
still on to write them back. Invalidate afterwards, so nothing stale is
left to be hit if the caches are enabled again. */
cache_clean_all();
__asm__ volatile("dsb sy" ::: "memory");
__asm__ volatile("mrc p15, 0, %0, c1, c0, 0" : "=r"(sctlr));
sctlr &= ~((1UL << 2) | (1UL << 12)); /* SCTLR.C, SCTLR.I */
__asm__ volatile("mcr p15, 0, %0, c1, c0, 0" : : "r"(sctlr) : "memory");
__asm__ volatile("dsb sy" ::: "memory");
__asm__ volatile("isb" ::: "memory");
cache_invalidate_dcache_all();
cache_invalidate_icache_all();
__asm__ volatile("dsb sy" ::: "memory");
__asm__ volatile("isb" ::: "memory");
}
@@ -50,6 +50,12 @@ void cache_clean_all(void);
void cache_enable(void);
/* Clean, then disable both caches, then invalidate. Needed by any measurement
that wants more than one cold pass in a single run: cache_enable is
one-way, and a second cold reading is otherwise impossible. */
void cache_disable_all(void);
unsigned int cache_enabled(void);
/* Decoded L1 data-cache geometry, from CCSIDR. Sizing a cache benchmark
@@ -193,6 +193,7 @@ unsigned int board_service_count;
#define BOARD_ISR_SECTION
#endif
__attribute__((aligned(64)))
static BOARD_ISR_SECTION void board_irq_service_body(unsigned long intid);
void board_irq_service(unsigned long intid)
@@ -212,6 +213,11 @@ void board_irq_service(unsigned long intid)
}
/* Aligned for the same reason as cache_workload: this body is timed, and
without a fixed alignment the figure moves with unrelated code changes
elsewhere in the image. */
__attribute__((aligned(64)))
static BOARD_ISR_SECTION void board_irq_service_body(unsigned long intid)
{
if (intid == GICV3_SPURIOUS_INTID)