Short answers to what readers ask most about this topic.
01How do you design a distributed job scheduler that runs each job exactly once?
Store every scheduled occurrence as a database row with a primary key of schedule name plus slot time, so only one instance can create it. Workers claim the row with FOR UPDATE SKIP LOCKED and an expiring lease. Because a worker can crash after the side effect but before recording success, you get at-least-once delivery and make the handler idempotent to get effectively-once results.
02Why does a cron job run twice when I scale to multiple instances?
Cron is a per-machine daemon with no knowledge of other hosts, so every instance carrying the same crontab fires independently. A job every 5 minutes on 3 instances produces 864 executions a day instead of 288. Rolling deploys that keep old and new containers alive together cause the same duplicates.
03What is the difference between an advisory lock and a lease?
A Postgres session-level advisory lock is held until it is released or the database session ends, so it depends on the connection staying alive. A lease carries an expiry timestamp that the holder must renew with heartbeats, so a crashed or stalled worker loses it automatically. Leases suit long jobs and recovery, while advisory locks suit cheap mutual exclusion such as no-overlap rules.
04What should a scheduler do about runs missed during downtime?
Choose per job between skipping missed slots, running once for the latest slot, or running every missed slot. With a 15-minute job and an outage from 10:02 to 10:52, three slots are missed, so those policies fire 0, 1 or 3 runs. Add a maximum lateness so very old slots are marked skipped rather than run.
05How long should retries wait between attempts?
Use exponential backoff with full jitter: wait a random time between zero and min(cap, base times 2 to the power attempt minus 1). With a 30 second base and 900 second cap, the ceilings are 30, 60, 120, 240, 480 and 900 seconds. Cap the attempts, send non-retryable errors straight to a dead state, and alert on dead runs.
Design a Distributed Job Scheduler That Runs Each Job Once
How to design a distributed job scheduler that runs each job once: slot rows, Postgres SKIP LOCKED leases, advisory locks, catch-up, retries and idempotent handlers.
To run each scheduled job once across several instances, give every time slot its own database row, let workers claim it with FOR UPDATE SKIP LOCKED and an expiring lease, and make handlers idempotent. Exactly-once is impossible end to end, so you get at-least-once delivery with effectively-once results. Add a catch-up policy, jittered retries and a lag alert.
The classic way to learn this is the day you add a second replica and the nightly report goes out twice. Cron was never wrong: it fired on time, on every machine that had the crontab, and nothing in it knows the other machines exist.
This is a worked design, not a war story, so every number below is derived from stated assumptions with the arithmetic shown. The behaviour of FOR UPDATE SKIP LOCKED and advisory locks comes from the PostgreSQL documentation, the lease and fencing reasoning from Martin Kleppmann, and the backoff maths from the AWS Builders Library. The stack is the one I reach for on a single VPS: Postgres, Node and NestJS.
Why does cron on several instances run every job twice?
Cron is a per-machine daemon. It reads its own crontab, watches its own clock and starts the command, with no shared state and no way to ask whether another host already did. Put the same schedule on three replicas, or let a rolling deploy keep the old and the new container alive together for a minute, and each one fires independently. Clock drift makes it worse, because the copies do not even fire at the same instant, so a naive lock taken just before the job can be released before the slower copy arrives.
Take a job every 5 minutes on 3 instances. A day has 24 times 60 divided by 5, which is 288 slots. Three instances fire each slot, so 3 times 288 gives 864 executions where you wanted 288, and 576 of them are duplicates. If the job sends an email or charges a card, that is 576 wrong side effects a day. The fix is not a smarter lock around the job; it is to stop treating the wall clock as the trigger and make the slot itself a record.
How do you make one scheduled slot fire only once?
Store the schedule and every concrete occurrence of it in the database. A schedule row says when the next slot is due. A run row represents one slot, and its primary key is the schedule name plus the slot time. Any instance may try to create that row, and the primary key guarantees only one succeeds. Below is the schema and the tick that every instance can run every few seconds.
-- A schedule says WHEN. A job_run is one concrete slot of it.
CREATE TABLE schedules (
name text PRIMARY KEY,
every interval NOT NULL, -- e.g. '5 minutes'
next_run_at timestamptz NOT NULL,
enabled boolean NOT NULL DEFAULT true
);
CREATE TABLE job_runs (
schedule_name text NOT NULL REFERENCES schedules(name),
scheduled_for timestamptz NOT NULL, -- the slot, not the wall clock
status text NOT NULL DEFAULT 'pending', -- pending | running | done | dead
attempts int NOT NULL DEFAULT 0,
run_after timestamptz NOT NULL DEFAULT now(),
locked_by text,
locked_until timestamptz,
last_error text,
finished_at timestamptz,
PRIMARY KEY (schedule_name, scheduled_for) -- this line is the whole "once per slot" guarantee
);
-- The tick. Every instance may run it every few seconds; they cannot collide.
WITH due AS (
SELECT name, next_run_at, every
FROM schedules
WHERE enabled AND next_run_at <= now()
ORDER BY next_run_at
LIMIT 50
FOR UPDATE SKIP LOCKED -- a second instance skips rows the first holds
), fired AS (
INSERT INTO job_runs (schedule_name, scheduled_for)
SELECT name, next_run_at FROM due
ON CONFLICT DO NOTHING -- belt and braces: the slot already exists
RETURNING schedule_name
)
UPDATE schedules s
SET next_run_at = d.next_run_at + d.every -- one slot at a time = "run every missed slot"
FROM due d
WHERE s.name = d.name;
Two things make this safe. First, the slot time comes from the stored next_run_at, never from each instance's own clock, and the only clock consulted is now() on the database, which is a single clock. Second, FOR UPDATE SKIP LOCKED means an instance that finds a due schedule already locked by another instance moves on instead of waiting. The PostgreSQL documentation notes that skipping locked rows gives an inconsistent view of the data, which makes it unsuitable for general queries but is exactly what you want for queue-like tables where each worker should take a different row.
How does claim-with-lease work with SKIP LOCKED?
Creating the slot row is not running the job. A worker must claim a run, and claiming has to survive the worker dying. A lease does that: instead of a lock that lives until someone releases it, the claim carries an expiry, locked_until. A run is claimable when it is pending and due, or when it is running but its lease has expired. Both cases go through the same query.
-- Claim one run with a lease. Pending work and work whose lease expired look the same.
UPDATE job_runs
SET status = 'running',
attempts = attempts + 1, -- doubles as the fencing token
locked_by = $1,
locked_until = now() + make_interval(secs => $2)
WHERE (schedule_name, scheduled_for) = (
SELECT schedule_name, scheduled_for
FROM job_runs
WHERE (status = 'pending' AND run_after <= now())
OR (status = 'running' AND locked_until < now()) -- the previous worker died or stalled
ORDER BY scheduled_for
LIMIT 1
FOR UPDATE SKIP LOCKED -- concurrent claimers skip, never wait
)
RETURNING schedule_name, scheduled_for, attempts;
Take a lease of 60 seconds, a heartbeat every 20 seconds and a worker that polls every 5 seconds, all of which are assumptions you tune. If a worker is killed mid-run, its lease expires within 60 seconds and the next poll finds the run, so the worst-case recovery delay is about 60 plus 5, or 65 seconds. The attempts column increments on every claim and doubles as a fencing token, which the retry section uses. For the queue mechanics behind this pattern, see my earlier post on the Postgres SKIP LOCKED job queue; this post is about deciding when jobs become due, not how they are consumed.
Set the heartbeat to about a third of the lease. One missed beat is then a blip, two are a warning, and three mean the worker is gone. A lease much shorter than a normal garbage-collection pause or network stall will reclaim runs that are still healthy.
Do you need advisory locks or a leader?
Often not, because the primary key already removes the race. But it helps to see the options side by side. Leaders and locks are tools for narrowing who ticks or for forbidding overlap, not a requirement for correctness once slots are rows.
Approach
Fires twice?
If the instance dies
Fits when
Cron on every instance
Yes, once per instance
Survivors keep firing
Only when the job is harmless to repeat
Cron on one designated instance
No
Nothing runs until a human notices
One box where downtime is acceptable
Advisory-lock leader
Not while the lock holds
Postgres frees the lock when the session ends and another instance takes over
A few schedules and Postgres already in the stack
Slot row plus claim with lease
Not per slot; a handler may still repeat
The lease expires and another worker reclaims the run
Several instances and a need for run history
A leader is a cheap optimisation: one instance runs the tick, the rest stay quiet. Postgres session-level advisory locks give you that in one call, since the PostgreSQL documentation says they are held until released explicitly or until the session ends. They also forbid overlap, because a long job can hold a lock keyed on a numeric schedule id and a second worker that cannot take it skips the run. Treat the lock as a hint and keep the primary key as the guarantee.
// Leader election lite: one dedicated connection holds a session-level advisory lock.
const LEADER_KEY = 421001; // any bigint your app agrees on; it names nothing in the database
const leader = await pool.connect(); // NOT pool.query: the lock lives on this session
const { rows } = await leader.query("SELECT pg_try_advisory_lock($1) AS got", [LEADER_KEY]);
if (rows[0].got) {
startTickLoop(); // only the leader materialises slots
} else {
leader.release(); // someone else leads; retry in a few seconds
}
// If this process dies, Postgres drops the session and the lock with it.
A session-level advisory lock belongs to one database connection, so take it on a dedicated connection and not through a pool that hands connections around, and check how a transaction-pooling proxy behaves before trusting it. Even then, a paused process can wake up believing it still leads after another instance has taken over, which is the failure Martin Kleppmann describes for lock leases. Never let leadership alone protect money.
Can a scheduler really run a job exactly once?
Not end to end, and it is better to say so than to promise it. A worker performs a side effect, then writes done. If it crashes between the two, the scheduler cannot tell whether the effect happened. It can re-run the job, which risks a duplicate, or drop it, which risks a loss. You choose at-least-once or at-most-once; exactly-once delivery is not on the menu.
What you can build is effectively-once: at-least-once delivery plus an idempotent handler, so a second run changes nothing. The slot gives you a perfect idempotency key, because the same slot always carries the same schedule name and scheduled time no matter how many attempts it takes. The attempts counter fences the scheduler's own rows, since a zombie worker whose lease was reclaimed matches zero rows when it tries to finish. As Kleppmann points out, a fencing token only protects a resource that checks it, so for the outside world you need the idempotency key.
// Wrong: a retry or a reclaimed lease sends the second invoice.
await db.query("INSERT INTO invoices (customer_id, amount) VALUES ($1, $2)", [id, amount]);
// Right: the slot is the idempotency key, so the second attempt is a no-op.
const periodKey = name + ":" + slot.toISOString(); // "monthly-billing:2026-11-01T00:00:00.000Z"
await db.query(
`INSERT INTO invoices (period_key, customer_id, amount)
VALUES ($1, $2, $3)
ON CONFLICT (period_key, customer_id) DO NOTHING`,
[periodKey, id, amount],
);
Pass the same key to anything you call. Payment and email providers commonly accept an idempotency key on a request, so give them the slot key as well, and keep your own unique constraint as the final backstop. My post on idempotency keys in API design covers how to build the receiving side.
What happens to runs missed while the scheduler was down?
Decide this per job, in advance, because the right answer differs. Take a job every 15 minutes and an outage from 10:02 to 10:52. The slots 10:15, 10:30 and 10:45 were due and missed, so 3 slots. You have three honest policies.
Skip: fire 0 runs and resume at 11:00. Right for a reminder that is pointless an hour late.
Run once: fire 1 run, labelled with the latest missed slot, 10:45. Right for a refresh or a sync where the newest state is all that matters.
Run all: fire 3 runs, one per missed slot. Right when each slot is a separate fact, such as a per-period close that must exist for every period.
-- Catch-up policy "run once": jump next_run_at past now() instead of one slot at a time.
SET next_run_at = d.next_run_at + d.every * (
1 + floor(extract(epoch FROM (now() - d.next_run_at)) / extract(epoch FROM d.every))
)
The tick above advances one slot at a time, which is run-all. The statement above switches it to run-once by jumping next_run_at past now(). To also express skip, add a maximum lateness and mark any slot older than it as skipped instead of pending. Whichever you pick, write it next to the schedule so that the next engineer does not rediscover it during an incident.
How should retries and backoff work?
Retry with exponential backoff and full jitter, and cap both the delay and the attempts. The AWS Builders Library describes full jitter as sleeping a random time between zero and the exponential ceiling, so that clients that failed together do not retry together. With a base of 30 seconds and a cap of 900 seconds, the ceilings for attempts 1 to 6 are 30, 60, 120, 240, 480 and then 960 capped to 900 seconds. Their sum is 1,830 seconds, about 30.5 minutes, and with full jitter the average wait is half of each ceiling, so roughly 15 minutes. After attempt 6 the run goes to dead for a human. Here is the worker loop with the lease, the heartbeat and the fenced writes.
import { Pool } from "pg";
const pool = new Pool();
const LEASE_SECONDS = 60;
const HEARTBEAT_MS = 20_000; // a third of the lease: two missed beats still leave margin
const MAX_ATTEMPTS = 6;
const BASE_SECONDS = 30;
const CAP_SECONDS = 900;
// Full jitter: wait a random time between 0 and the exponential ceiling.
function backoffSeconds(attempt: number): number {
const ceiling = Math.min(CAP_SECONDS, BASE_SECONDS * 2 ** (attempt - 1));
return Math.random() * ceiling;
}
// Every write is fenced on attempts: a worker whose lease was taken over matches 0 rows.
const FINISH_SQL = `UPDATE job_runs SET status = 'done', finished_at = now(), locked_until = NULL
WHERE schedule_name = $1 AND scheduled_for = $2 AND attempts = $3`;
const HEARTBEAT_SQL = `UPDATE job_runs SET locked_until = now() + make_interval(secs => $4)
WHERE schedule_name = $1 AND scheduled_for = $2 AND attempts = $3 AND status = 'running'`;
const FAIL_SQL = `UPDATE job_runs
SET status = CASE WHEN attempts >= $4 THEN 'dead' ELSE 'pending' END,
run_after = now() + make_interval(secs => $5),
last_error = $6, locked_until = NULL
WHERE schedule_name = $1 AND scheduled_for = $2 AND attempts = $3`;
async function runOnce(workerId: string, handler: (slot: Date) => Promise<void>) {
const { rows } = await pool.query(CLAIM_SQL, [workerId, LEASE_SECONDS]);
if (rows.length === 0) return;
const { schedule_name: name, scheduled_for: slot, attempts } = rows[0];
const beat = setInterval(
() => pool.query(HEARTBEAT_SQL, [name, slot, attempts, LEASE_SECONDS]),
HEARTBEAT_MS,
);
try {
await handler(slot); // must be safe to run twice: see the handler section
await pool.query(FINISH_SQL, [name, slot, attempts]);
} catch (err) {
await pool.query(FAIL_SQL, [name, slot, attempts, MAX_ATTEMPTS,
backoffSeconds(attempts), String(err)]);
} finally {
clearInterval(beat);
}
}
Two details matter. Separate errors that retrying cannot fix, such as a validation failure, and send those to dead at once instead of burning six attempts. And make sure the retry window fits the schedule: a 30-minute retry window on a 5-minute job overlaps six later slots, so either forbid overlap with a per-schedule lock or accept that the job may run concurrently with its own retry.
What should you monitor on a job scheduler?
A scheduler fails quietly: nothing crashes, a job simply stops happening. So alert on the absence of success, not only on errors. Three numbers per schedule cover most of it, and one query produces them.
-- Three numbers per schedule that tell you the scheduler is healthy.
SELECT schedule_name,
now() - max(finished_at) FILTER (WHERE status = 'done') AS since_last_success,
count(*) FILTER (WHERE status = 'dead') AS dead_runs,
now() - min(scheduled_for) FILTER (WHERE status = 'pending') AS oldest_pending_age
FROM job_runs
GROUP BY schedule_name;
Alert when since_last_success exceeds two intervals, for example 10 minutes on a 5-minute job. This catches a stopped scheduler, which error alerts never see.
Alert when dead_runs is above zero. A dead run is a job that exhausted its attempts and needs a person.
Watch oldest_pending_age. A steadily rising value means workers are too few or too slow, even though every run eventually succeeds.
Log one structured line per attempt with schedule name, slot, attempt number and worker id, so one slot can be followed across retries and workers.
Test the failure on purpose: kill a worker mid-run and confirm another finishes the slot within the lease plus the poll interval, about 65 seconds with the numbers above.
Do not page on a single failed attempt; retries exist to absorb those. Page on the three signals above, and keep the attempt logs for the post-mortem.
A scheduler should never rely on the clock to decide who runs a job. Make each slot a row with a unique key, claim it with a lease, assume it will sometimes run twice and make the handler harmless when it does. Then decide your catch-up policy, jitter your retries and alert on silence.