Short answers to what readers ask most about this topic.
01What is the difference between pub/sub and a message queue?
A message queue delivers each message to exactly one consumer from a pool, so it distributes work. Pub/sub delivers each message to every subscriber, so it broadcasts an event to many independent readers. Queues answer who does this work, pub/sub answers who needs to know.
02Is Redis Pub/Sub a message queue?
No. The Redis documentation describes Pub/Sub as at-most-once delivery: if a subscriber is disconnected or fails, the message is lost for good. It also stores nothing. For queue-like behaviour use a job library such as BullMQ, or Redis Streams if you need persisted messages.
03Can a message queue do pub/sub?
Some brokers can, by giving each subscriber its own queue bound to the same exchange or topic. Log-based systems such as Kafka do it with consumer groups: instances inside one group share messages like a queue, and each additional group receives all of them. So the two models are less separate than they first look.
04When should I use pub/sub instead of a queue?
Use pub/sub when one event must reach several independent services, such as a completed order triggering loyalty, analytics and a dashboard. Use a queue when exactly one worker should perform a task, such as sending a receipt. If subscribers cannot afford to miss events, choose a durable log rather than transient pub/sub.
05What is a dead-letter queue and do I need one?
A dead-letter queue holds messages that failed every retry so they can be inspected instead of blocking or looping forever. You need a dead-letter path whenever messages can be poison, for example malformed payloads. Some brokers provide it, while NATS JetStream publishes a max-deliveries advisory and expects you to park the message yourself.
Pub/Sub vs Message Queue: Differences and When to Use Each
Pub/sub vs message queue: competing consumers versus fan-out, delivery guarantees, ordering, replay, consumer groups, and a checklist for jobs, notifications and events.
A message queue hands each message to one consumer, so workers share a backlog of jobs. Pub/sub delivers each message to every subscriber, so many services react to one event. Queues add acknowledgements and retries, transient pub/sub adds neither, and log-based brokers such as Kafka and NATS JetStream use consumer groups to offer both.
The question usually arrives as a design review comment: should this be a queue or pub/sub? In a POS or carwash ERP the same event, a finished wash, has to produce a WhatsApp receipt, a loyalty point, a refreshed display board and a row in a report. Those are not the same kind of work, and a broker that is right for one is wrong for another.
This post is the conceptual decision guide, not a tool tutorial. It compares the two models, shows the smallest real snippet for BullMQ, Redis Pub/Sub and a Kafka consumer group, derives what each failure mode costs with arithmetic, and ends with a checklist. Every behaviour claimed is taken from the official documentation listed at the end.
What is the core difference between pub/sub and a message queue?
The difference is who receives a message. In a queue, a pool of consumers reads from the same backlog and each message goes to exactly one of them, which is point-to-point delivery with competing consumers. Adding workers makes the backlog drain faster. In pub/sub, the publisher does not know who is listening: the Redis documentation describes publishers that are not programmed to send to specific receivers, and subscribers that receive what they asked for without knowing who published. Adding subscribers adds readers of the same stream, not throughput.
So the two answer different questions. A queue answers: who will do this piece of work? Pub/sub answers: who needs to know that this happened? Work that must happen once, such as charging a card or sending a receipt, belongs in a queue. A fact that several unrelated services react to, such as a wash being completed, belongs in pub/sub.
How does a message queue with competing consumers work?
The producer puts a job on a named queue and any worker attached to that queue may take it. In BullMQ the producer calls add on a Queue, and workers are created with the same queue name. Starting the worker in two processes gives you two competing consumers, and each job reaches only one of them.
import { Queue, Worker } from "bullmq";
const connection = { host: "127.0.0.1", port: 6379 };
// Producer: one job per receipt. attempts + backoff are the retry policy.
const receipts = new Queue("receipts", { connection });
await receipts.add(
"send-receipt",
{ washId: 8841, channel: "whatsapp" },
{ attempts: 3, backoff: { type: "exponential", delay: 1000 } },
);
// Consumer: start this in two processes with the SAME queue name.
// Each job is handed to exactly one of them (competing consumers).
const worker = new Worker(
"receipts",
async (job) => {
await sendReceipt(job.data); // throw = failed attempt, BullMQ retries it
},
{ connection },
);
// After the last attempt the job lands in the failed set. Look at it, alert on it.
worker.on("failed", (job, err) => alertOps(job?.id, err.message));
The part that matters is what happens when a worker fails. BullMQ takes a retry policy as job options: attempts sets the total number of attempts and backoff sets a fixed or exponential delay in milliseconds. When a worker cannot renew its lock on a job, the job is marked stalled and moved back to waiting so another worker processes it again. That is the queue promise in one sentence: a job is not forgotten because one worker died, and the price is that a job can run more than once.
How does pub/sub fan-out work, and what does it not remember?
A subscriber declares interest in a channel and the broker pushes every message on that channel to every subscriber currently connected. In Redis that is SUBSCRIBE and PUBLISH. The documentation notes that subscribers receive messages in the order they were published, and that a subscribed client in RESP2 can only issue subscription commands, so real applications keep a separate connection for publishing.
import Redis from "ioredis";
// A connection in subscribed mode cannot run normal commands (RESP2),
// so the subscriber and the publisher are two separate connections.
const sub = new Redis();
const pub = new Redis();
// Every process that runs this gets EVERY message. No group, no queue.
await sub.subscribe("wash.completed");
sub.on("message", (channel, payload) => {
refreshDisplayBoard(JSON.parse(payload));
});
// PUBLISH returns how many subscribers it reached. 0 means the message is gone.
const reached = await pub.publish("wash.completed", JSON.stringify({ washId: 8841 }));
What Redis Pub/Sub does not do is remember. The documentation states at-most-once delivery semantics: a message is delivered once if at all, and if the subscriber hits an error or a network disconnect the message is forever lost. Nothing is stored for a subscriber that is not connected. For stored messages the same page points to Redis Streams, which persist messages and support both at-most-once and at-least-once delivery.
Worked example: if a subscriber restarts for 10 seconds while a channel carries 2 messages per second, it never sees 10 x 2 = 20 messages, and nothing reports it. That is fine for a display board that redraws on the next event, and a silent bug for anything that must be counted.
Delivery, ordering, retention: how do queues and pub/sub compare?
Three properties decide most designs: what happens when a consumer is down, whether order is kept, and whether you can read a message again. The table compares a classic job queue, transient pub/sub (Redis Pub/Sub, and core NATS which is also at-most-once), and a log-based broker that uses consumer groups (Kafka, NATS JetStream).
Question
Job queue (BullMQ)
Transient pub/sub (Redis Pub/Sub)
Log with consumer groups (Kafka)
Who gets a message
One worker out of the pool
Every connected subscriber
One consumer per group, every group
Delivery guarantee
A stalled or failed job runs again, so at-least-once in effect
At-most-once: lost if the subscriber is down or errors
At-least-once when you commit the offset after processing
Consumer offline
Job waits in the queue until a worker returns
Messages published meanwhile are gone
Records wait in the log, the group resumes from its committed offset
Replay old messages
No, a finished job is consumed
No, nothing is stored
Yes, a consumer can specify an earlier offset
Ordering
Picked up roughly in order, but completion order is not guaranteed with several workers or retries
Delivered in the order published
Kept within a partition, not across partitions
Failure handling
attempts, backoff and a failed set
None, it is the subscriber's problem
Re-read from the last committed offset, or nak and AckWait redelivery in JetStream
Read the table by rows, not by columns. If you need the row Consumer offline to say nothing is lost, transient pub/sub is out. If you need Replay old messages, a classic job queue is out. The cell on ordering is the one people skip: with several workers and retries, jobs finish in a different order than they started, so a design that needs strict order has to partition by key.
How do consumer groups combine queue and pub/sub in Kafka and NATS?
A log-based broker appends messages to a durable log and keeps a position, called an offset, per consumer group. Confluent's design documentation says each partition is consumed by exactly one consumer within each consumer group at any given time. Put several instances in one group and they share the partitions like competing queue workers. Create a second group and it receives the same records independently, which is fan-out. One mechanism, both models.
import { Kafka } from "kafkajs";
const kafka = new Kafka({ clientId: "qilap-api", brokers: ["localhost:9092"] });
// Same topic, two groups. Within a group the partitions are shared out
// (queue behaviour); across groups every group sees everything (pub/sub).
const loyalty = kafka.consumer({ groupId: "loyalty" });
const analytics = kafka.consumer({ groupId: "analytics" });
await loyalty.connect();
await loyalty.subscribe({ topics: ["wash.completed"], fromBeginning: true });
await loyalty.run({
eachMessage: async ({ topic, partition, message }) => {
await addPoints(JSON.parse(message.value!.toString()));
// offset is committed after this resolves: a crash re-delivers (at-least-once)
},
});
// analytics.run(...) is identical with its own groupId, and its own offsets.
The cost is that the broker stores everything until retention removes it, and that a group that crashes restarts at its last committed offset, so it may reprocess records it had already handled. Confluent documents exactly that: after reassignment the initial position is the last committed offset. A consumer can also specify an earlier offset to reconsume data, which is the replay row in the table. The KafkaJS option fromBeginning: true is how a brand new group asks for the whole history.
A new subscriber added to a log-based topic can backfill history by starting at an earlier offset, which a transient pub/sub channel cannot do at all. If a future service might need yesterday's events, that alone can justify the log.
What happens when a consumer fails: acknowledgement, redelivery and dead letters?
An acknowledgement is the consumer telling the broker the work is done. Without one the broker must assume failure and deliver again. NATS JetStream documents four responses: ack for success, nak for redeliver, term for stop trying, and in-progress to reset the AckWait timer. A plain nak redelivers immediately, which can become a tight retry loop, so give it a delay as in the snippet.
// JetStream pull consumer named "shipping", explicit acks.
const c = await js.consumers.get("ORDERS", "shipping");
const messages = await c.consume();
for await (const m of messages) {
try {
await handle(m);
m.ack(); // work done: server will not redeliver
} catch {
m.nak(10_000); // redeliver after 10 s instead of in a tight loop
}
}
Worked example with the documented defaults: AckWait defaults to 30 seconds and MaxDeliver defaults to -1, meaning unlimited. A poison message whose handler hangs is redelivered every 30 seconds forever. Setting MaxDeliver to 5 caps it at 5 x 30 = 150 seconds of waiting before the server stops delivering it.
Stopping is not the same as handling. JetStream has no built-in dead-letter queue, and publishes a max-deliveries advisory instead, so you subscribe to it and park the message yourself. BullMQ keeps a job in its failed set after the last attempt, depending on your removal settings. Either way a dead-letter path is something you wire up and then watch, otherwise failures simply pile up unseen.
Which should you use for notifications, background jobs and events?
Map each piece of work to the model that matches what it needs. The mapping below comes from the carwash ERP example, where one wash completion triggers several unrelated reactions.
Notifications such as a WhatsApp or email receipt: a queue. It must happen once, retry on failure and survive a worker restart.
Background jobs such as report generation or image resizing: a queue. You want a backlog, more workers when it grows and a failed set to inspect.
Events such as wash completed: pub/sub, and a log with consumer groups if any subscriber must not miss one or a new service may need history. Plain Redis Pub/Sub is for live, disposable signals like refreshing a screen.
When you are unsure, ask these four questions in order. The first no or yes that decides it usually decides the whole design.
Is it work that exactly one worker should do? Use a queue.
Do several independent services each need every event? Use pub/sub, with one consumer group or subscription per service.
Is losing a message while a consumer is down unacceptable? Rule out transient pub/sub and pick a queue or a log.
Might a future consumer need old events? Pick a log-based broker with retention.
Choose by what happens to a message after it is published. If one worker must do it, use a queue. If many services must hear about it, use pub/sub. If many services must hear about it reliably and a late joiner may want the past, use a log with consumer groups. Whichever you pick, make the handler idempotent, because retries and redelivery mean a message can arrive twice.