Skip to content
  • Overview
  • Curriculum
    • FLUMental maths and numerical fluency
    • MKTMarkets and products
    • CSData structures and algorithms
    • PYPython and data for quants
    • NUMNumerical methods
    • SYSSystems and low latency
      • 1Computer architecture

        • Architecture: caches, branch prediction and SIMD
      • 2C++ for trading

        • C++ for trading: RAII, moves and what belongs on the hot path
      • 3Concurrency

        • Concurrency: atomics, memory ordering and lock-free queues
      • 4Networking

        • Networking: multicast market data, kernel bypass and gap recovery
      • 5Measurement

        • Latency, the memory hierarchy and why the tail is the number
      • 6Order book engineering

        • Order book engineering: data structures and the operations they serve

Practise

  • Question bank
  • Mental arithmetic
  • Market simulator
  • Arbitrage trees
  • Horse racing
  • Bid book
  • Screening tests
  • Mock papers

Reference

  • Formula reference
  • Search

Your record

  • Review queue
  • Progress
  • Leaderboard
  • Profile
  • Invite friends
AccountSend feedback
  1. Curriculum
  2. /Quantitative development
  3. /Systems and low latency
  4. /Measurement

Latency, the memory hierarchy and why the tail is the number

SYS · Chapter 5·12 min read·Asked at Hudson River Trading, Jump, Citadel Securities, Optiver

After this lesson you should be able to

  • Quote the latency numbers every low-latency interview assumes you know.
  • Explain why a mean latency is close to useless and a percentile is not.
  • Name the hot-path practices that follow from the memory hierarchy.

A trading system is judged on the slowest requests it serves, not the typical one, because the busy microsecond is exactly when the opportunity and the risk both arrive. That single observation drives most of what low-latency engineering looks like, from preallocation to cache layout to how you benchmark.

OperationOrder of magnitude
L1 cache reference~1 ns
Branch mispredict~5 ns
L2 cache reference~4 ns
Main memory reference~100 ns
Same-datacentre round trip~500 µs
NVMe read~100 µs
Context switch~1–5 µs
System call~100 ns–1 µs
Table 5.1 · The numbers to know. The one ratio that matters: main memory is roughly a hundred times slower than L1. A cache miss on the hot path costs more than the arithmetic it was fetching data for.

Proposition 5.2

Percentiles, not averages

Report p50, p99, p99.9 and the maximum. A mean latency conflates the common path with the garbage-collection pause, the page fault and the reallocation, and it hides precisely the events that cost money. A system with a 2 μs2\,\mu s2μs mean and a 2 ms2\,ms2ms p99.9 is worse, for trading, than one with a 5 μs5\,\mu s5μs mean and a 20 μs20\,\mu s20μs p99.9.

Holds when

  • Jitter — the spread of the distribution — is often the target, not the mean.
  • Percentiles do not compose: the p99 of a pipeline is not the sum of its stages’ p99s.
  • Measure tick-to-trade end to end, since stage-by-stage numbers miss the queueing between them.
509999.990450900Tick-to-tradePercentileLatency, microseconds
Figure 5.3 · The average is not the number. A median of four microseconds and a thousandth-percentile of two hundred and ten. The tail is not noise around the average — it is a different mechanism, usually a page fault, a garbage collection or a lock, and it turns up exactly when the market is busiest.

Why the tail lands at the worst moment. Slow paths are not randomly distributed in time. Buffers fill when message rates spike; allocators run out of pooled memory when activity is highest; caches are evicted when more symbols are active. So the p99.9 latency occurs disproportionately during the bursts that follow news — which is exactly when quotes are stale and adverse selection is worst. The tail is correlated with the opportunity, which is why it is the number a trading firm optimises.

What follows for hot-path code

  • Preallocate. No allocation, no resizing, no locks on the path.
  • Keep data contiguous and hot: arrays of structs the loop actually reads, not pointer chases.
  • Avoid false sharing — two threads writing different variables on one 64-byte cache line serialise.
  • Prefer branch-free or predictable branches; a mispredict costs more than the work it skipped.
  • Do the work early: precompute anything that does not depend on the incoming message.

The rest of this lesson is in Premium

You have read the opening. 11 more sections follow, including 4 worked examples and 3 quick checks.

Start the free 7-day trialSign in

Nothing is charged for 7 days, and you can cancel before then. Or read Complexity: reading it off, and deriving it in full, free.

← Networking: multicast market data, kernel bypass and gap recoveryOrder book engineering: data structures and the operations they serve →
On this page
  • The numbers to know
  • Percentiles, not averages
  • The average is not the number

QuantMax · 141 lessons · 1342 questions · c5c0caa

  • Premium
  • Arbitrage trees
  • Horse racing
  • Invite friends
  • Account
  • About QuantMax
  • Terms
  • Privacy

Firm names identify publicly reported question patterns and nothing more. QuantMax is not affiliated with, endorsed by, or recruiting for any firm named in the curriculum. Everything you do in lessons and the question bank is kept to your account.