System Design

Designing a Notification System at Scale

How to design a notification system for push, email and SMS at scale: requirements, architecture, queues, templates, preferences, retries and deduplication.

Phone, email and message icons connected to a central notification service
Illustration: Backend Architect / AI-generated.

Key takeaways

  • Accept notifications through an API, validate them, then hand off to queues per channel.
  • Workers call push, email and SMS providers with retries, rate limits and deduplication.
  • Respect user preferences and quiet hours, and track delivery status.
On this page

Almost every product sends notifications: order updates, password resets, reminders, marketing, and messages from a chat system to users who are offline. A notification system looks simple until it has to send millions of messages across several channels without losing, duplicating or spamming any of them.

1. Requirements

Functional: send push, email and SMS; templates with variables; user preferences and opt-outs; scheduled sends; delivery tracking.

Non-functional: reliable (no lost critical messages), scalable to bursts, low latency for time-sensitive messages such as login codes, and no duplicates.

2. High-level architecture

  1. Notification API receives requests from other services, authenticates and validates them.
  2. Preference and template service checks opt-outs and quiet hours and renders the message.
  3. Queues per channel and priority buffer work; see Kafka vs RabbitMQ vs SQS.
  4. Channel workers send through providers: push gateways, email providers and SMS aggregators.
  5. Status store records sent, delivered, failed and opened events.
  6. Scheduler triggers future and recurring notifications.

3. Priorities

Separate queues for transactional (password resets, security alerts) and bulk (newsletters, promotions) traffic, so a marketing blast never delays a login code.

4. Reliability

  • Retries with exponential backoff for provider failures.
  • Dead-letter queues for messages that keep failing, so they can be inspected.
  • Provider failover: switch to a backup email or SMS provider if the primary degrades.
  • Idempotency: attach a unique ID to every notification so retries and at-least-once delivery don’t create duplicates. See idempotency keys.

5. Rate limiting

Providers limit throughput, and users shouldn’t be flooded. Apply rate limiting per provider, per user and per notification type, for example no more than a few marketing pushes per day.

6. User preferences

Store per-user, per-channel, per-category settings, plus time zone and quiet hours. Check them before rendering, and honour unsubscribes immediately; email and SMS rules in many countries require it.

7. Scaling

Workers are stateless and scale horizontally with queue depth. Cache templates and preferences (see caching strategies), and shard the status store by user or notification ID as it grows.

8. Monitoring

Track send rate, provider errors, delivery and open rates, queue lag and end-to-end latency for critical messages, with alerts when they drift.

Frequently asked questions

How do you avoid sending duplicate notifications?

Assign each notification a unique ID, record processed IDs and make workers idempotent.

Should notifications be synchronous?

No. Accept the request quickly, queue it and send asynchronously.

How do you handle time zones?

Store each user’s time zone and schedule sends in local time, respecting quiet hours.

Sources

  1. Google — Site Reliability Engineering book

Every article is edited by a human and checked against our editorial policy. Spotted a mistake? Tell us.

Keep reading