Short answers to what readers ask most about this topic.
01How do you design a scalable notification system?
Separate the event from the delivery. Business code emits an event, an ingester writes a deduplicated notification row, and a worker per channel (push, email, SMS) sends it after checking preferences. Add retries with backoff, a dead-letter queue, a rate limit per provider and webhook-based status tracking.
02Why should push, email and SMS have separate queues?
Each channel has its own provider, its own rate limit and its own failure modes. With one shared queue, a throttled SMS provider can hold back push messages that are ready to send. Separate queues let you tune concurrency and retries per channel and see each backlog on its own.
03How do you prevent duplicate notifications?
Build a dedup key from the intent, such as event id, user and channel, and put a unique constraint on it. Insert with ON CONFLICT DO NOTHING and send only if a row was inserted. Do not use a random id per attempt, because every retry would then look like a new message.
04How should a notification system handle retries and failures?
Classify the result first. Permanent errors such as a 400 or a hard bounce go straight to a dead-letter queue, while transient errors such as a 429 or 500 are retried with exponential backoff and jitter. Honour Retry-After when the provider sends it and set a maximum age, for example one hour.
05How do you track whether a notification was actually delivered?
Store the provider's message id when you send, then update the row from the provider's webhooks. SES publishes events such as Delivery and Bounce, and Twilio posts each status change to a callback URL. Callbacks can arrive out of order, so only let the status move forward.
How to Design a Scalable Notification System (Push, Email, SMS)
A notification system is an event ingester, one queue per channel, a preference check, a dedup constraint, retries into a dead-letter queue and webhook-driven delivery tracking.
Design a scalable notification system by turning every business event into one deduplicated database row, fanning it out to a separate queue per channel (push, email, SMS), checking preferences and suppressions before sending, retrying transient failures with backoff into a dead-letter queue, throttling each provider, and updating status from provider webhooks.
The first notification feature in most systems is one function: after the order is paid, call an email API. It works until the day the provider is slow, the request times out, the job retries, and a customer receives the same receipt three times.
This post is the architecture that sits between a business event and a phone buzzing. It stays on a single stack I work with, Postgres and Redis on a modest server, and every figure below is arithmetic from the stated assumptions or taken from a provider document listed at the end, not a benchmark. Queue mechanics and outbound webhooks have their own posts here, so this one covers what is specific to notifications.
What does a notification system actually have to do?
It has to turn a fact (order 8841 was paid) into zero or more messages on the right channels, once each, in the user's language, respecting their choices. The business code should emit only the event. Everything about channels, templates and providers belongs behind a boundary, so that adding WhatsApp later changes one worker and not forty call sites. The two payload shapes below are the contract.
type Channel = "push" | "email" | "sms";
type Priority = "critical" | "bulk"; // OTP and receipts must never queue behind a campaign
// What the business code emits. It knows nothing about channels or providers.
interface NotificationEvent {
eventId: string; // stable id from the source system, e.g. "order-8841-paid"
type: "order.paid" | "otp.requested" | "invoice.overdue";
userId: string;
occurredAt: string; // ISO 8601
data: Record<string, string | number>;
}
// What a channel worker consumes. One job per (notification, channel).
interface SendJob {
notificationId: string; // primary key of the notifications row
dedupKey: string; // eventId + ":" + userId + ":" + channel
channel: Channel;
priority: Priority;
userId: string;
templateKey: string; // "order.paid"
templateVersion: number; // pinned at enqueue time, so a retry renders the same copy
locale: "en" | "id";
vars: Record<string, string | number>;
attempt: number; // 0 on first try
notBefore?: string; // set by the quiet-hours rule or by a backoff
}
Two decisions are baked into that shape. First, the template is referenced by key and version rather than rendered at enqueue time, so a retry an hour later produces the same copy, and a rendering error (SES reports one as a Rendering Failure event) is caught in the worker instead of poisoning the producer. Second, the notification row is written in the same database transaction as the business change, so an event that committed can never be lost between the commit and the queue.
Why use one queue per channel instead of one shared queue?
Because the three channels fail in different ways, at different speeds, under different limits. A shared queue lets a rate-limited SMS provider hold back push messages that were ready to go. Separate queues give each channel its own concurrency, its own retry policy and its own backlog you can read at a glance, and a priority split inside each one keeps OTP codes ahead of campaigns.
Channel
Where status comes from
What failure looks like
Retry rule worth encoding
Push (FCM)
The response to the send request
400, 401, 403, 404 are permanent; 429 and 500 are transient
Abort on 4xx; honour Retry-After on 429, default 60 s; exponential backoff on 500; wait at least 10 s
Email (SES)
Events published to SNS: Send, Delivery, Bounce, Complaint, Reject, DeliveryDelay
A Permanent bounce means the address is dead; a Complaint means the user flagged you
Never retry a Permanent bounce; add the address to suppressions at once
SMS (Twilio)
A StatusCallback POST per status change
Status ends at delivered, undelivered or failed, with an ErrorCode
Callbacks can arrive out of order, so never let a status move backwards
The table is the reason the retry code later in this post returns a typed outcome rather than throwing. A worker that cannot tell a permanent failure from a transient one will retry a dead address eight times and train your provider to distrust your sender reputation.
How do you handle user preferences and opt-outs?
Keep preferences and suppressions in two separate tables, because they have different authors. A preference is something the user chose, per category and per channel. A suppression is something a provider told you, such as a hard bounce, a complaint or a STOP reply, and it overrides every preference because sending anyway damages your standing. The schema below makes a missing preference row mean the default, so a new category never needs a backfill.
CREATE TYPE channel AS ENUM ('push', 'email', 'sms');
-- One row per (user, category, channel). A missing row means "use the default",
-- so a new category never needs a backfill across every user.
CREATE TABLE notification_preferences (
user_id uuid NOT NULL,
category text NOT NULL, -- 'transactional', 'security', 'marketing'
channel channel NOT NULL,
enabled boolean NOT NULL DEFAULT true,
quiet_start time, -- local time, NULL = no quiet hours
quiet_end time,
timezone text NOT NULL DEFAULT 'Asia/Jakarta',
updated_at timestamptz NOT NULL DEFAULT now(),
PRIMARY KEY (user_id, category, channel)
);
-- Written by webhooks, NOT by the user. A hard bounce or an unsubscribe lands here
-- and overrides any preference row, because the provider has already told you no.
CREATE TABLE channel_suppressions (
channel channel NOT NULL,
address text NOT NULL, -- email, E.164 phone number or push token
reason text NOT NULL, -- 'hard_bounce', 'complaint', 'unsubscribe', 'stop'
created_at timestamptz NOT NULL DEFAULT now(),
PRIMARY KEY (channel, address)
);
For email, make the unsubscribe a single request. RFC 8058 defines the List-Unsubscribe-Post header with the value List-Unsubscribe=One-Click so a mail client can unsubscribe the user with one POST, and your handler should write straight to channel_suppressions. Check the preference at the moment of sending, not at enqueue time, because a user who opts out while a job waits in the queue expects it to stop. Security and transactional categories should not be switchable off, and the UI should say so.
How do you stop the same notification being sent twice?
Give every intended message a deterministic dedup key and let a database unique constraint decide who is first. The key is built from things that identify the intent, not the attempt: the source event id, the user and the channel, for example order-8841-paid:user-17:email. A random UUID per attempt, the usual mistake, makes every retry look like a new message.
CREATE TABLE notifications (
id uuid PRIMARY KEY DEFAULT gen_random_uuid(),
dedup_key text NOT NULL UNIQUE, -- the whole mechanism is this constraint
user_id uuid NOT NULL,
channel channel NOT NULL,
status text NOT NULL DEFAULT 'queued',
provider_id text, -- the provider's message id, for webhooks
created_at timestamptz NOT NULL DEFAULT now()
);
-- Wrong: SELECT to check, then INSERT. Two consumers can both see "not there".
-- Right: let the unique index decide, atomically.
INSERT INTO notifications (dedup_key, user_id, channel)
VALUES ($1, $2, $3)
ON CONFLICT (dedup_key) DO NOTHING
RETURNING id;
-- 1 row back -> first time we have seen this event: enqueue the job.
-- 0 rows back -> a replay or a duplicate delivery: stop, send nothing.
The INSERT with ON CONFLICT DO NOTHING is atomic: PostgreSQL documents that RETURNING yields only the rows actually inserted, so zero rows means somebody already owns this key. That covers a producer that replays an event, and a queue that delivers a job twice. It does not cover the gap after the provider accepted the message but before you recorded it, which is why the next section matters.
Exactly-once sending is not available when the provider call and your database are two systems. If the process dies after the provider accepted a message and before you store its id, the retry will send again. Where the provider accepts an idempotency key, pass your dedup key; where it does not, accept an occasional duplicate and prefer it over a lost OTP.
How should retries and the dead-letter queue work?
Classify the outcome first, then decide. A permanent error goes straight to the dead-letter queue, because retrying a 400 only repeats the 400. A transient error is retried with exponential backoff and jitter. FCM's own guidance is to wait at least 10 seconds, to honour Retry-After on a 429, defaulting to 60 seconds, to use exponential backoff on 500s, and to stop retrying after 60 minutes.
const BASE_MS = 10_000; // never retry sooner than 10 s
const MAX_ATTEMPTS = 8; // 10+20+40+80+160+320+640+1280 = 2,550 s of ceilings
const DEADLINE_MS = 60 * 60_000; // give up after an hour, the provider's own advice
// Full jitter: wait a random time between 0 and the exponential ceiling, so a
// thousand jobs that failed together do not all come back together.
const backoffMs = (attempt: number) =>
Math.random() * BASE_MS * 2 ** attempt;
type Outcome =
| { kind: "sent"; providerId: string }
| { kind: "retry"; afterMs?: number } // 429 with Retry-After, 5xx, timeout
| { kind: "dead"; reason: string }; // 400, 401, 403, 404: retrying cannot help
async function handle(job: SendJob, firstSeenAt: number) {
const out = await sendViaProvider(job);
if (out.kind === "sent") return markSent(job, out.providerId);
const exhausted =
out.kind === "dead" ||
job.attempt + 1 >= MAX_ATTEMPTS ||
Date.now() - firstSeenAt > DEADLINE_MS;
if (exhausted) return moveToDeadLetter(job, out); // keep the payload, add the reason
// The provider's Retry-After beats our own schedule when it sends one.
const delay = out.afterMs ?? backoffMs(job.attempt);
return requeue({ ...job, attempt: job.attempt + 1 }, delay);
}
Worked example with a 10 second base and doubling: the ceilings are 10, 20, 40, 80, 160, 320, 640 and 1,280 seconds, which sum to 2,550 seconds, about 42.5 minutes. A ninth attempt would add 2,560 seconds and end at 5,110, past the 3,600 second hour, so eight attempts is the natural cap. The dead-letter queue is not a graveyard: store the payload, the last error and the attempt count, alert on its depth, and give yourself a one-click replay for the day a provider has a bad afternoon.
How do you respect provider rate limits and fail over?
Put a token bucket in front of every provider call, sized to the limit on your own account, and make the worker wait instead of fail when the bucket is empty. FCM, for instance, throttles with 429 once a quota bucket is exhausted, and its guidance is to ramp traffic up over at least a 60 second window rather than start at full rate. Failover is a second provider behind a circuit breaker, used only for provider-side trouble such as timeouts and 5xx, never for a bad recipient.
// A token bucket per provider. In-memory is enough on ONE worker process;
// with several workers the bucket must live in Redis or the limit multiplies.
class TokenBucket {
private tokens: number;
private last = Date.now();
constructor(private ratePerSec: number, private burst: number) {
this.tokens = burst;
}
tryTake(): boolean {
const now = Date.now();
this.tokens = Math.min(
this.burst,
this.tokens + ((now - this.last) / 1000) * this.ratePerSec,
);
this.last = now;
if (this.tokens < 1) return false;
this.tokens -= 1;
return true;
}
}
const smsPrimary = new TokenBucket(20, 20); // 20/s is an example, read your own account limit
async function sendSms(job: SendJob): Promise<Outcome> {
if (!smsPrimary.tryTake()) return { kind: "retry", afterMs: 50 }; // wait for a token
if (breaker.isOpen("sms-primary")) return sendViaSecondary(job); // failover
const out = await callPrimary(job);
// Fail over on PROVIDER trouble only. A 400 for a bad number fails identically
// on the secondary, and you would pay twice to learn the same thing.
if (out.kind === "retry") breaker.recordFailure("sms-primary");
return out;
}
Arithmetic shows why priority lanes matter. Suppose an SMS account allows 20 messages a second, an assumption for illustration. A campaign of 12,000 messages takes 12,000 divided by 20, which is 600 seconds, ten minutes, and an OTP queued behind it waits up to that long. With a separate critical lane that drains first, the same OTP waits for at most a few tokens. Failover has its own cost: a timeout is ambiguous, since the first provider may have delivered, so a failed-over send can duplicate.
Run the bucket in Redis once you have more than one worker process. An in-memory bucket per worker silently multiplies your limit by the worker count, and the first sign is a wave of 429s from the provider.
How do you track delivery with provider webhooks?
A send call only tells you the provider accepted the message. Delivery, bounce and failure arrive later as webhooks: SES publishes Send, Delivery, Bounce, Complaint and more to SNS, and Twilio posts each status change to your StatusCallback URL. Look the notification up by the provider's message id, which is why the notifications table stores it. Twilio states plainly that callbacks are not guaranteed to arrive in the order they were sent.
-- Rank the statuses, then only ever move forward. The webhook for "sent" can
-- arrive AFTER the one for "delivered"; without the guard it would overwrite it.
UPDATE notifications
SET status = $2
WHERE provider_id = $1
AND (CASE status
WHEN 'queued' THEN 1
WHEN 'sent' THEN 2
WHEN 'delivered' THEN 3
WHEN 'undelivered' THEN 3
WHEN 'failed' THEN 3
ELSE 0
END)
< (CASE $2
WHEN 'sent' THEN 2
WHEN 'delivered' THEN 3
WHEN 'undelivered' THEN 3
WHEN 'failed' THEN 3
ELSE 0
END);
-- Terminal bad news also feeds the suppression table, so the next send is skipped:
-- email Bounce with bounceType Permanent -> reason 'hard_bounce'
-- email Complaint -> reason 'complaint'
So the update must only move forward. Rank the statuses and write the new one only if its rank is higher, as in the query above, otherwise a late sent overwrites delivered and your dashboard lies. Make the webhook handler idempotent too, since providers retry their own callbacks, and let terminal bad news write to the suppression table so the next send is skipped without anyone looking.
What is the checklist before shipping a notification system?
Use this list as a review gate. Each item maps to a failure described above, and an unchecked item is a known incident waiting for a date.
Business code emits events only; one ingest boundary maps events to channels and templates.
The notification row commits in the same transaction as the business change.
A unique dedup key, built from intent rather than attempt, guards every insert.
One queue per channel with a critical lane, and a preference plus suppression check at send time.
Outcomes are typed; permanent errors go to the dead-letter queue, transient ones back off with jitter to a cap.
A token bucket per provider, a circuit breaker for failover, and forward-only status updates from webhooks.
If you can only do three on day one, do the dedup key, the preference check and the retry classification. Those three prevent the visible mistakes, namely duplicates, unwanted messages and hammering a dead address, and everything else can be added behind the same contract.
A scalable notification system is mostly a set of refusals: refuse to send twice, refuse to send to someone who said no, refuse to retry what cannot succeed, and refuse to trust the order of webhooks. Build the boundary and the constraints first and the provider choices become swappable details.