Latency, the memory hierarchy and why the tail is the number
SYS · Chapter 512 min readAsked at Hudson River Trading, Jump, Citadel Securities, Optiver
After this lesson you should be able to
- Quote the latency numbers every low-latency interview assumes you know.
- Explain why a mean latency is close to useless and a percentile is not.
- Name the hot-path practices that follow from the memory hierarchy.
A trading system is judged on the slowest requests it serves, not the typical one, because the busy microsecond is exactly when the opportunity and the risk both arrive. That single observation drives most of what low-latency engineering looks like, from preallocation to cache layout to how you benchmark.
| Operation | Order of magnitude |
|---|---|
| L1 cache reference | ~1 ns |
| Branch mispredict | ~5 ns |
| L2 cache reference | ~4 ns |
| Main memory reference | ~100 ns |
| Same-datacentre round trip | ~500 µs |
| NVMe read | ~100 µs |
| Context switch | ~1–5 µs |
| System call | ~100 ns–1 µs |
Proposition 5.2
Percentiles, not averages
Report p50, p99, p99.9 and the maximum. A mean latency conflates the common path with the garbage-collection pause, the page fault and the reallocation, and it hides precisely the events that cost money. A system with a mean and a p99.9 is worse, for trading, than one with a mean and a p99.9.
Holds when
- Jitter — the spread of the distribution — is often the target, not the mean.
- Percentiles do not compose: the p99 of a pipeline is not the sum of its stages’ p99s.
- Measure tick-to-trade end to end, since stage-by-stage numbers miss the queueing between them.
Why the tail lands at the worst moment. Slow paths are not randomly distributed in time. Buffers fill when message rates spike; allocators run out of pooled memory when activity is highest; caches are evicted when more symbols are active. So the p99.9 latency occurs disproportionately during the bursts that follow news — which is exactly when quotes are stale and adverse selection is worst. The tail is correlated with the opportunity, which is why it is the number a trading firm optimises.
What follows for hot-path code
- Preallocate. No allocation, no resizing, no locks on the path.
- Keep data contiguous and hot: arrays of structs the loop actually reads, not pointer chases.
- Avoid false sharing — two threads writing different variables on one 64-byte cache line serialise.
- Prefer branch-free or predictable branches; a mispredict costs more than the work it skipped.
- Do the work early: precompute anything that does not depend on the incoming message.
The rest of this lesson is in Premium
You have read the opening. 11 more sections follow, including 4 worked examples and 3 quick checks.
Nothing is charged for 7 days, and you can cancel before then. Or read Complexity: reading it off, and deriving it in full, free.