aytugyayman

RAM · Benchmark

What does a memory benchmark measure?

Aytug Yayman 2 min read

Building a memory benchmark showed me why one ns or GB/s figure is not enough. Its meaning depends on where the data lives and how the test accesses it.

Separate latency from bandwidth

Latency is measured through dependent address reads: the next address comes from the previous read. This makes it harder to hide waiting behind many independent requests. Bandwidth uses separate read, write and copy streams.

ZenStation has a PageLocal path that favours accesses within pages and a FullRandom path that moves across them. The latter can expose more address-translation overhead. Results from these access patterns should not be compared as if they came from the same test.

Avoid measuring cache as RAM

A working set that fits in cache may describe cache behaviour rather than DRAM. Separate L1, L2, L3 and RAM results make the distinction visible. ZenTimings v1.44 presents cache and RAM bandwidth and latency in one table.

ZenStation considers the selected buffer size, available memory and CPU topology together. Buffer size is especially important on CPUs with large L3 caches. Two-CCD systems need separate L3 domains; I do not assume every core shares one L3.

Keep comparable records

Core placement, warm-up runs, background activity and large-page use can affect a result. The application records its core choice and options with the result. Failure to use large pages is reported rather than hidden.

When evaluating a configuration change, I repeat the same profile, buffer size and access pattern. I compare the distribution of runs, the previous record and the memory timings, rather than relying on one particularly good run. A change in measurement-method version also requires reviewing comparability with old results.

A performance result does not by itself establish system stability. This article explains the measurement method; an example result is not a target for every memory kit.

Sources and related work

All articles