About Us Contact Us Write for Us Advertise
Home > System Design > High-Level Design > Design a Rate Limiter — System Design Interview
System Design › High-Level Design

Design a Rate Limiter — System Design Interview

Design a distributed rate limiter end to end: where it lives (the API gateway), which algorithm, and the core challenge — sharing counters across many servers with Redis and atomic increments.

Shiv Pandey
Shiv Pandey
Oct 08, 2026 | 3 views
Design a Rate Limiter — System Design Interview

🟠 High-Level Design · Senior interview favourite

"Design a rate limiter" is a favourite because it's a real component every large API needs, and it forces you to reason about distributed state. You've likely met the rate-limiting algorithms already; this is the system design around them — where the limiter lives, how it shares counters across many servers, and what happens when it's overwhelmed. Let's design it properly, whiteboard-style.

The whole design in one line: at the API gateway, for each incoming request identify the caller, check a shared counter in Redis, and either let it through or reject it with 429 — all in under a millisecond so the check itself doesn't slow the API.

Step 1 — Requirements

  • Functional: cap requests per client (e.g. 100/min per user or API key); reject excess with HTTP 429 Too Many Requests and a Retry-After; limits must be configurable per rule.
  • Non-functional: low latency (the check runs on every request, so it must be sub-millisecond), accurate across many servers, highly available, and it must fail open — if the limiter breaks, don't block all traffic.

Step 2 — Where does it live?

🟠 Put the limiter at the edge — in the API gateway or a reverse proxy, before requests reach your application. Two reasons: bad traffic is rejected early (your app servers never waste work on it), and the logic lives in one shared place instead of being copy-pasted into every service.

clients API Gateway+ rate limiter Redisshared counters App servers check+incr if allowed over limit → 429

Step 3 — Pick the algorithm

🟠 From the algorithms, token bucket is the usual pick: it caps the average rate but allows short bursts, and it stores just two numbers per client (tokens + last-refill time) — tiny and fast. Sliding-window counter is the other strong choice when you want smoother enforcement without boundary bursts. Either is cheap enough to run on every request.

Step 4 — The core challenge: shared state

🔴 Here's the whole reason this is a system-design question. Your limit is "100/min per user," but you run many gateway instances behind a load balancer. If each instance keeps its own in-memory counter, a user spread across 10 instances gets 10× the limit. The counter must be shared and centralised — and the standard answer is Redis: in-memory (sub-millisecond), and every gateway reads/writes the same key.

🔴 But concurrency bites: if two requests read the same counter at once, both see "99" and both allow — a race that lets you exceed the limit. The fix is to make check-and-increment atomic. Two clean ways:

  • Redis INCR with an expiry: INCR is atomic and returns the new count; set the key to expire at the window's end. Simple, and perfect for fixed/sliding-window counting.
  • A small Lua script: Redis runs it atomically server-side, so you can do token-bucket refill + deduct in one indivisible step. Ideal for token bucket.
key = "rl:" + userId + ":" + currentMinute
count = INCR key            # atomic, returns new value
if count == 1:  EXPIRE key 60
if count > 100: reject → 429
else:           allow

"Shared counter in Redis, updated atomically via INCR-with-expiry or a Lua script" is the sentence that wins this question.

Step 5 — Response & headers

When a request is over the limit, return 429 with a Retry-After telling the client how long to wait. It's also good practice to send X-RateLimit-Limit, X-RateLimit-Remaining, and X-RateLimit-Reset so well-behaved clients can self-throttle instead of blindly retrying.

Step 6 — The senior edge cases

🔴 What separates a great answer:

  • Fail open: if Redis is down, do you block all traffic? Usually no — you let requests through (fail open) so a limiter outage doesn't become a full outage. State this trade-off explicitly.
  • Latency of the Redis hop: a network call per request adds latency. Mitigate by co-locating Redis, or using a local approximate counter synced periodically to Redis (fast, slightly less exact) for extreme scale.
  • What to key on: user ID for logged-in users, API key for partners, IP for anonymous traffic (with care behind proxies). Often several rules layered together.
  • Redis as a bottleneck/SPOF: run it replicated/clustered so it's not a single point of failure, and shard keys across nodes for throughput.

The interview-ready walkthrough

"I'd run the limiter at the API gateway so bad traffic dies early. I'd use a token bucket (average cap + small bursts, two numbers per user) and store counters in Redis so the limit is global across all gateway instances. The critical detail is making check-and-increment atomic — an INCR with expiry or a Lua script — to avoid races. Over-limit gets 429 + Retry-After. I'd fail open if Redis is down, run Redis clustered so it's not a SPOF, and key limits on user/API-key/IP." That covers placement, algorithm, the distributed problem, and the edge cases — the full arc."

What to read next

← Design a URL shortener · Design a news feed →

Related Articles

Design a URL Shortener (TinyURL / bit.ly) — System Design
System Design › High-Level Design

Design a URL Shortener (TinyURL / bit.ly) — System Design