Latency, Throughput & Availability: How We Measure Fast and Reliable
Latency, throughput and availability explained from scratch — what each measures, the highway analogy, why averages lie (p50 vs p99 tail latency), the availability nines table, and how these numbers drive every architecture decision.
🟢 Fresher-friendly → 🟠 senior nuance
In a system design interview, someone will eventually ask "how fast is it?" or "how reliable does it need to be?" — and vague words like "pretty fast" won't cut it. Engineers measure these things precisely with three words: latency, throughput, and availability. Get comfortable with them and you'll instantly sound like you know what you're doing. This guide explains each from scratch, shows how they trade off against each other, and gives you the real-world numbers worth memorising.
Latency — how long one request takes
Latency is the time between asking for something and getting it, usually measured in milliseconds (ms). Click a link and the page shows in 200 ms? That's the latency. Lower is better. Humans start noticing sluggishness around 100 ms and feel real friction past ~1 second.
Analogy: latency is how long you personally wait in the coffee queue — from joining the line to holding your cup.
Latency has many sources that add up: the network trip (physics — data can't beat the speed of light, so distance matters), time spent waiting in queues on busy servers, database lookups, and disk reads. Cutting latency is why we put servers near users (CDNs), keep answers in memory (caches), and avoid slow disk when we can.
Throughput — how much you handle per second
Throughput is how many requests your system can process per unit of time — often written as RPS (requests per second) or QPS (queries per second). Higher is better. A system might handle 10,000 requests per second; that's its throughput.
Analogy: throughput is how many customers the whole coffee shop serves per hour — across all baristas and tills.
Here's the key insight beginners miss: latency and throughput are not the same thing, and improving one can hurt the other. A single super-fast barista gives low latency but low throughput. Add ten baristas and throughput soars — but if the shop gets so crowded that people wait to even reach a till, individual latency can get worse. Big systems fight to keep both good at once.
The road analogy that makes it click
🔵 Picture a highway:
- Latency = how long it takes one car to drive from A to B (the speed limit and distance).
- Throughput = how many cars pass a point per minute (the number of lanes).
Adding lanes (more servers) raises throughput but doesn't make any single car faster. Raising the speed limit (faster code, closer servers) lowers latency. And when too many cars pile on, you get a traffic jam — throughput collapses and everyone's latency spikes. That jam is exactly what happens to an overloaded server, and it's why we add load balancers and queues.
The trap: averages lie (meet p99)
🟠 If your average latency is 50 ms, is that good? Maybe not. Averages hide the unlucky users. That's why engineers talk in percentiles:
- p50 (median): half of requests are faster than this.
- p99: 99% of requests are faster than this — so the slowest 1% are worse.
Why obsess over the slowest 1%? Because on a page that makes 100 back-end calls, there's a very good chance at least one hits that slow p99 — so the whole page feels slow to almost everyone. This is called tail latency, and taming it is a genuinely senior-level skill. When an interviewer asks "what's your latency target?", answering "p99 under 200 ms" instead of "fast" signals real experience.
Availability — how often you're up at all
Availability is the percentage of time your system is working and reachable. It's measured in "nines," and every extra nine is dramatically harder (and more expensive) to achieve:
| Availability | Nickname | Downtime per year |
|---|---|---|
| 99% | "two nines" | ~3.65 days |
| 99.9% | "three nines" | ~8.8 hours |
| 99.99% | "four nines" | ~52 minutes |
| 99.999% | "five nines" | ~5 minutes |
Going from 99% to 99.99% isn't a small polish — it means cutting yearly downtime from days to minutes, which usually requires removing every single point of failure (redundant servers, replicated databases, multiple data centres). That's why a hobby project happily lives at "two nines" while a bank chases "five." How we buy those nines — replication, failover, redundancy — runs through the whole scalability and replication story.
The three, side by side
How they show up in an interview
🟠 Near the start of any design, a strong candidate states target numbers out loud — it frames every decision that follows:
"Let's assume 10M daily users, ~5,000 requests/sec at peak (throughput), a target of p99 under 200 ms (latency), and 99.9% availability. That means the design needs multiple app servers, caching, and a replicated database so no single failure takes us down."
See how the numbers drive the architecture? High throughput → many servers + a load balancer. Low latency → caching + nearby servers. High availability → redundancy + replication. You're not guessing; you're reasoning from targets. That's the whole game, and the step-by-step version is in how to approach a system design interview.
What to read next
- Scaling from Zero to Millions — how we actually hit these targets
- How to Approach a System Design Interview — the framework that uses these numbers
- Scalability: Vertical vs Horizontal — the idea under throughput
- ← The complete System Design guide (hub)
← Client, server, database & API · Scaling from zero to millions →
Frequently Asked Questions
What is the difference between latency and throughput?
Latency is how long a single request takes (measured in milliseconds, lower is better). Throughput is how many requests you handle per second (measured in RPS, higher is better). Adding servers raises throughput but does not make one request faster; that is latency.
What is p99 latency and why does it matter?
p99 latency is the value 99% of requests are faster than — so it captures the slowest 1%. It matters because a page making many back-end calls will likely hit at least one slow call, so tail latency makes the whole page feel slow to almost everyone.
What do the availability nines mean?
They describe uptime: 99% (two nines) is about 3.65 days of downtime per year, 99.9% is about 8.8 hours, 99.99% is about 52 minutes, and 99.999% is about 5 minutes. Each extra nine is far harder and costlier to achieve.