About Us Contact Us Write for Us Advertise
Home > System Design > High-Level Design > Design a Notification System — System Design Interview
System Design › High-Level Design

Design a Notification System — System Design Interview

Design a multi-channel notification system (push, SMS, email) at scale: the queue-based async architecture, per-channel workers, third-party providers, retries and dead-letter queues, dedupe, rate limiting, and priority.

Shiv Pandey
Shiv Pandey
Oct 10, 2026 | 4 views
Design a Notification System — System Design Interview

🟠 High-Level Design · Senior interview favourite

"Design a notification system" — the service that sends push notifications, SMS, and emails at scale (order confirmations, "someone liked your photo," OTPs). It looks like a simple "send a message" task, but doing it for millions of users across multiple channels, reliably, without spamming anyone, is a proper distributed-systems problem. It's also a beautiful showcase for queues and the fan-out-to-workers pattern. Let's design it.

The whole design in one line: a service accepts "notify user X about Y," drops it on a queue, and channel-specific workers pull from the queue and hand off to the right provider (APNs/FCM for push, Twilio for SMS, SendGrid for email) — with preferences, retries, and deduplication along the way.

Step 1 — Requirements

  • Functional: send notifications over multiple channels (push, SMS, email, in-app); support many event types; respect user preferences (opt-outs, channel choice, quiet hours); use templates.
  • Non-functional: huge scale and spiky load; highly available; reliable (an OTP must arrive); low latency for time-sensitive ones; and it must not spam (dedupe, rate-limit).

Step 2 — The central insight: decouple with a queue

🟠 The single most important design choice: never send notifications synchronously from the code that triggered them. Sending involves slow, unreliable third-party providers; if "place order" waited on the SMS API, a slow provider would stall checkout. Instead, the trigger just publishes a message to a queue and returns immediately. Background workers do the actual sending. This decoupling gives you buffering for spikes, retries on failure, and independent scaling — the whole reason a queue sits at the heart of this design.

Step 3 — Architecture

Services(triggers) NotificationService Queue(Kafka) Push worker SMS worker Email worker APNs / FCM Twilio SendGrid checks prefs, dedupe, rate-limit, templating

The flow: a service calls the Notification Service, which validates the request, checks the user's preferences (are they opted in to this on this channel?), applies a template, and drops a message per channel onto the queue. Channel workers (push, SMS, email) consume their queue and call the relevant third-party provider. Separate workers per channel means you scale and fail independently — a down SMS provider never blocks push.

Step 4 — The details that make it robust

🔴 The senior substance lives in these concerns:

  • Retries & dead-letter queue: providers fail transiently. Workers retry with backoff; messages that keep failing go to a dead-letter queue for inspection instead of being lost or retried forever.
  • Idempotency / deduplication: at-least-once queue delivery means a worker might process a message twice — and double-texting a user is bad. Give each notification an idempotency key and dedupe so it's sent once.
  • Rate limiting: cap how many notifications a user gets (per hour/day) so you don't spam them, and respect provider rate limits.
  • Preferences & opt-out: a preference service is the source of truth for "does this user want this?" — checked before sending. Unsubscribes are legally required for email.
  • Priority: an OTP is urgent; a marketing blast isn't. Separate high- and low-priority queues so critical notifications aren't stuck behind a bulk campaign.

Step 5 — Data model

notifications:  id, user_id, type, channel, template_id,
                status, created_at, idempotency_key
preferences:    user_id, channel, event_type, enabled
templates:      template_id, channel, subject, body
device_tokens:  user_id, platform, push_token   # for APNs/FCM

You also log each notification's status (queued → sent → delivered → failed) for analytics, debugging, and "why didn't my OTP arrive?" support.

Step 6 — Scaling

The queue absorbs spikes (a big campaign becomes a backlog workers drain steadily, not a crash). Each channel's workers scale horizontally by adding consumers. Because it's all async and decoupled, you can push huge volume without the triggering services ever slowing down — exactly what a queue-based architecture buys you.

The interview-ready walkthrough

"I'd make it fully asynchronous: triggering services call a Notification Service that checks user preferences, applies a template, and publishes to a queue. Per-channel workers (push/SMS/email) consume and call providers like APNs, Twilio, SendGrid — so channels scale and fail independently. For reliability I'd add retries with backoff and a dead-letter queue, idempotency keys to dedupe at-least-once deliveries, rate limits to avoid spamming, and priority queues so OTPs jump ahead of marketing. The queue absorbs spikes and keeps the triggering services fast." That hits decoupling, reliability, and anti-spam — the whole checklist."

What to read next

← Design a chat system · Design typeahead autocomplete →

Related Articles

Design a Rate Limiter — System Design Interview
System Design › High-Level Design

Design a Rate Limiter — System Design Interview

Design a Chat System (WhatsApp / Messenger) — System Design
System Design › High-Level Design

Design a Chat System (WhatsApp / Messenger) — System Design

Design a News Feed (Facebook / Twitter) — System Design
System Design › High-Level Design

Design a News Feed (Facebook / Twitter) — System Design

Design a URL Shortener (TinyURL / bit.ly) — System Design
System Design › High-Level Design

Design a URL Shortener (TinyURL / bit.ly) — System Design