Design a Rate Limiter — System Design Interview
Design a distributed rate limiter end to end: where it lives (the API gateway), which algorithm, and the core challenge — sharing counters across many servers with Redis and atomic increments.
🟠 High-Level Design · Senior interview favourite
"Design a rate limiter" is a favourite because it's a real component every large API needs, and it forces you to reason about distributed state. You've likely met the rate-limiting algorithms already; this is the system design around them — where the limiter lives, how it shares counters across many servers, and what happens when it's overwhelmed. Let's design it properly, whiteboard-style.
429 — all in under a millisecond so the check itself doesn't slow the API.
Step 1 — Requirements
- Functional: cap requests per client (e.g. 100/min per user or API key); reject excess with HTTP 429 Too Many Requests and a
Retry-After; limits must be configurable per rule. - Non-functional: low latency (the check runs on every request, so it must be sub-millisecond), accurate across many servers, highly available, and it must fail open — if the limiter breaks, don't block all traffic.
Step 2 — Where does it live?
🟠 Put the limiter at the edge — in the API gateway or a reverse proxy, before requests reach your application. Two reasons: bad traffic is rejected early (your app servers never waste work on it), and the logic lives in one shared place instead of being copy-pasted into every service.
Step 3 — Pick the algorithm
🟠 From the algorithms, token bucket is the usual pick: it caps the average rate but allows short bursts, and it stores just two numbers per client (tokens + last-refill time) — tiny and fast. Sliding-window counter is the other strong choice when you want smoother enforcement without boundary bursts. Either is cheap enough to run on every request.
Step 4 — The core challenge: shared state
🔴 Here's the whole reason this is a system-design question. Your limit is "100/min per user," but you run many gateway instances behind a load balancer. If each instance keeps its own in-memory counter, a user spread across 10 instances gets 10× the limit. The counter must be shared and centralised — and the standard answer is Redis: in-memory (sub-millisecond), and every gateway reads/writes the same key.
🔴 But concurrency bites: if two requests read the same counter at once, both see "99" and both allow — a race that lets you exceed the limit. The fix is to make check-and-increment atomic. Two clean ways:
- Redis
INCRwith an expiry:INCRis atomic and returns the new count; set the key to expire at the window's end. Simple, and perfect for fixed/sliding-window counting. - A small Lua script: Redis runs it atomically server-side, so you can do token-bucket refill + deduct in one indivisible step. Ideal for token bucket.
key = "rl:" + userId + ":" + currentMinute count = INCR key # atomic, returns new value if count == 1: EXPIRE key 60 if count > 100: reject → 429 else: allow
"Shared counter in Redis, updated atomically via INCR-with-expiry or a Lua script" is the sentence that wins this question.
Step 5 — Response & headers
When a request is over the limit, return 429 with a Retry-After telling the client how long to wait. It's also good practice to send X-RateLimit-Limit, X-RateLimit-Remaining, and X-RateLimit-Reset so well-behaved clients can self-throttle instead of blindly retrying.
Step 6 — The senior edge cases
🔴 What separates a great answer:
- Fail open: if Redis is down, do you block all traffic? Usually no — you let requests through (fail open) so a limiter outage doesn't become a full outage. State this trade-off explicitly.
- Latency of the Redis hop: a network call per request adds latency. Mitigate by co-locating Redis, or using a local approximate counter synced periodically to Redis (fast, slightly less exact) for extreme scale.
- What to key on: user ID for logged-in users, API key for partners, IP for anonymous traffic (with care behind proxies). Often several rules layered together.
- Redis as a bottleneck/SPOF: run it replicated/clustered so it's not a single point of failure, and shard keys across nodes for throughput.
The interview-ready walkthrough
"I'd run the limiter at the API gateway so bad traffic dies early. I'd use a token bucket (average cap + small bursts, two numbers per user) and store counters in Redis so the limit is global across all gateway instances. The critical detail is making check-and-increment atomic — an INCR with expiry or a Lua script — to avoid races. Over-limit gets 429 + Retry-After. I'd fail open if Redis is down, run Redis clustered so it's not a SPOF, and key limits on user/API-key/IP." That covers placement, algorithm, the distributed problem, and the edge cases — the full arc."
What to read next
- Rate-Limiting Algorithms — token bucket, leaky bucket & the window methods in depth
- Design a URL Shortener — the other classic warm-up
- Caching Strategies — Redis, the store behind the counters
- ← The complete System Design guide (hub)