Skip to main content
Latency vs Throughputlesson 2 of 4 · 3 min read

Percentiles, Not Averages

Latency is not a bell curve

Latency does not follow a bell curve. Most requests cluster around a fast typical value. A long tail straggles out to many times that, dragged there by pauses for memory cleanup, cache misses, waiting on locks, and packets sent twice.

Averaging that produces a number describing almost nobody. It sits above the middle that most people experience and far below what the unlucky ones suffer.

What a percentile says

Percentiles describe somebody real, so your numbers describe somebody real. p50 is the median, the typical experience. p95 and p99 describe the slow tail, so a p99 of 900 milliseconds means one request in a hundred took longer than that. Watch p999 too on a serious service. At 10,000 requests a second, the slowest tenth of a percent is still ten people every second, all staring at a spinner.

Why the tail matters more than its share

Two effects make that tail matter more than its share of traffic suggests. Heavy users meet it most. Somebody making 50 requests in one session has roughly a 40 percent chance of catching at least one p99 event, and those are the customers you least want to annoy.

Fan-out multiplies the exposure. If rendering one page calls 30 services at once, the page waits for the slowest of them, so a single page load samples the tail 30 times. So large companies set their targets on p99, not averages. It is also why a backend with a fine median and an ugly tail wrecks the frontend built on it.

Percentiles cannot be averaged across servers or time windows once you start fixing this. The mean of ten per-machine p99s is not your fleet's p99, and it is usually optimistic. Combine the raw distributions instead, or use tooling that merges histograms properly.

the shape of it
p50: 90 mshalf are fasterp95: 400 ms5 in 100 slowerp99: 4 s1 in 100 slowerAverage: 99 msdescribes nobodyhides this
step 1 of 2
One hundred requests, slowest last. The average sits near the fast crowd and says nothing about the person waiting 4 seconds.

Worked example

Wei's team at a travel site has a search API averaging 120 ms, well inside its 200 ms target, yet support keeps hearing that search feels slow. He plots the distribution: p50 is 85 ms, p95 is 400 ms, and p99 is 2.3 seconds, driven by an occasional cold shard in Redis, the in-memory store behind the cache, and by pauses while the runtime reclaims memory on two old hosts. The average of 120 ms was arithmetic camouflage. At 600 requests per second, that p99 means six searches every second take over two seconds. Worse, the results page fans out to 8 downstream calls, so about one page in fifteen catches a slow one. After pinning the GC settings and warming the shard, p99 drops to 310 ms, and the complaint volume falls by half while the average barely moves.