Short answers to what readers ask most about this topic.
01What is the difference between two-phase commit and saga?
Two-phase commit coordinates one atomic commit across participants: all prepare and hold their locks, then all commit or all roll back. A saga commits each step locally and, if a later step fails, runs compensating actions to undo the earlier ones. 2PC gives isolation but can block; a saga stays available but exposes intermediate state.
02Does PostgreSQL support two-phase commit?
Yes, through PREPARE TRANSACTION, COMMIT PREPARED and ROLLBACK PREPARED, and any session can finish a prepared transaction. It is disabled by default because max_prepared_transactions defaults to 0, and changing it needs a server restart. The docs say the feature is meant for an external transaction manager, not for ordinary application code.
03Why is two-phase commit a blocking protocol?
After a participant votes yes it cannot commit or roll back on its own, because the other participants may have been told the opposite. It must wait for the coordinator's decision while still holding its locks. If the coordinator fails permanently, those participants stay in doubt and the locked rows stay unwritable.
04Can a saga guarantee isolation?
No. Each step commits immediately, so other transactions can see intermediate state before a compensation runs. You reduce the damage with semantic locks such as a PENDING status, commutative updates and rereading values before writing, but you cannot get the isolation of a single database transaction.
05When should I use a saga instead of two-phase commit?
Use a saga when steps cross services, vendors or HTTP APIs that cannot do prepare, such as a payment gateway, or when the process runs long enough that holding locks is unacceptable. Use 2PC only when you control every participant, they all support prepare and transactions are short. If all the data fits in one database, a normal transaction beats both.
Two-Phase Commit vs Saga: Distributed Transactions Compared
Two-phase commit vs saga for distributed transactions: how 2PC blocks and holds locks, what PostgreSQL PREPARE TRANSACTION offers, and when a saga fits better.
Use two-phase commit when every participant supports prepare, you operate them all, and transactions are short: it gives all-or-nothing atomicity but blocks and holds locks if the coordinator dies. Use a saga when steps cross services or vendors: each step commits locally and failures trigger compensating actions, trading isolation for availability.
A carwash ERP like Qilap takes a payment, deducts shampoo and wax from stock and marks the wash ticket as paid. On one Postgres database that is a single BEGIN and a single COMMIT, and nobody has to think about it. The moment the payment lives in a gateway and the stock lives in another service, the same sentence becomes a distributed transaction, and the real question is what happens when one half succeeds and the other does not.
This post compares the two standard answers, two-phase commit and the saga pattern, using a three-step checkout as the worked example. It leans on the PostgreSQL documentation, the Wikipedia article on two-phase commit, microservices.io and the original Sagas paper. The NestJS implementation of a saga is in my earlier post; this one is about choosing between them.
How does two-phase commit work?
A coordinator splits one distributed commit into two phases. In the prepare phase it asks every participant whether it can commit; each one makes its changes durable, keeps its locks and votes yes or no. In the commit phase, if every vote was yes, the coordinator tells all participants to commit; if any vote was no or never arrived, it tells them all to roll back.
// Two-phase commit, coordinator side. Three participants: orders-db, stock-db, payments-ledger.
// Phase 1: prepare. Each participant makes its work durable, KEEPS its locks, and votes.
votes = [orders.prepare(txid), stock.prepare(txid), payments.prepare(txid)] // 3 requests + 3 votes
log.write(txid, allYes(votes) ? "COMMIT" : "ABORT") // the decision must survive a coordinator crash
// Phase 2: the decision is delivered. Locks are released only when this arrives.
for (p of participants) p.finish(txid, decision) // 3 requests + 3 acks
Worked example with three participants, an orders database, a stock database and a payments ledger. Prepare costs 3 requests and 3 votes, the decision costs 3 requests and 3 acknowledgements, so one checkout is 3 + 3 + 3 + 3 = 12 messages, plus the coordinator writing its decision to a log between the phases. The number matters less than the gap: a participant holds its locks from the moment it votes yes until the decision arrives, and that gap is the cost of the whole protocol.
What happens when the coordinator crashes in two-phase commit?
Two-phase commit is a blocking protocol. Wikipedia calls this its greatest disadvantage: a participant that has voted yes must wait for the coordinator, and if the coordinator fails permanently, some participants never resolve their transactions.
A participant in that state cannot decide alone, because the others may have been told the opposite. So every row it locked stays locked. If the coordinator is down for 5 minutes, every row touched by the in-doubt checkouts is unwritable for 5 x 60 = 300 seconds, and every other order that wants the same stock row queues behind it. One dead process becomes a lock queue across several databases, which is the real price of 2PC.
A prepared transaction survives restarts on purpose, so nothing times it out for you. Someone or something must always finish it with a commit or a rollback.
Can PostgreSQL do two-phase commit?
Yes, through PREPARE TRANSACTION, COMMIT PREPARED and ROLLBACK PREPARED. The docs say a prepared transaction is fully stored on disk and can be committed from any session, not only the one that created it. The feature is off by default: max_prepared_transactions defaults to 0 and can only be set at server start, so enabling it needs a restart.
-- postgresql.conf (restart required, the default is 0 = feature disabled)
-- max_prepared_transactions = 10
-- Session 1: the stock database's half of the checkout
BEGIN;
UPDATE stock SET qty = qty - 2 WHERE sku = 'WAX-500' AND qty >= 2;
PREPARE TRANSACTION 'checkout-42-stock'; -- no longer tied to this session; state is on disk
-- Any session, even after a crash and restart: what is in doubt right now?
SELECT gid, prepared, owner, database FROM pg_prepared_xacts;
-- The coordinator saw every vote come back yes, so it finishes (from ANY session):
COMMIT PREPARED 'checkout-42-stock';
-- ...or, if any other participant voted no:
-- ROLLBACK PREPARED 'checkout-42-stock';
-- Until one of those runs, the UPDATE above still holds its row lock on WAX-500.
The same documentation page is blunt that PREPARE TRANSACTION is not intended for applications. It exists for an external transaction manager, and X/Open XA is the standard it names as an example. It also warns that a long-lived prepared transaction keeps its locks, interferes with VACUUM, and in extreme cases can lead to a shutdown to prevent transaction ID wraparound. Without a transaction manager closing them out promptly, it recommends leaving max_prepared_transactions at zero.
If you do enable it, the docs suggest a value of at least max_connections so every session can have one prepared transaction pending. My own addition: alert on any row in pg_prepared_xacts that is older than a minute, because a healthy coordinator finishes within milliseconds.
What is a saga, and how do choreography and orchestration differ?
A saga replaces one distributed transaction with a sequence of local transactions, each committed immediately in its own service. If a step fails on a business rule, the saga runs compensating transactions that undo the earlier steps, as microservices.io describes it. The idea comes from the 1987 Sagas paper by Garcia-Molina and Salem, written for long lived transactions that would otherwise hold locks for hours or days. There are two ways to coordinate the steps.
Choreography: each local transaction publishes a domain event that triggers the next local transaction in another service. There is no central component, but the flow is spread across services and harder to read.
Orchestration: an orchestrator tells each participant which local transaction to run. The order, retries and failure path live in one place, at the price of one more component to run.
For a three-step checkout I would pick orchestration, because the failure path is the part that needs reading and testing. Choreography earns its keep with two or three services and a flow that rarely changes. The NestJS implementation is in my earlier post: Saga pattern for distributed transactions in NestJS
How do compensating actions work in a saga?
Each step carries its own undo, and the undo is a new forward transaction, not a rollback. It cannot erase that a payment was authorised; it can only void or refund it. That is why compensations must be idempotent and retried until they succeed: the outage that failed the step may still be going on. The idempotency keys post covers the mechanics. The sketch below keeps the list of completed steps and unwinds it in reverse.
type Step<C> = {
name: string;
run: (ctx: C) => Promise<void>;
compensate: (ctx: C) => Promise<void>; // a NEW forward action: void, refund, release. Not a rollback.
};
const checkout: Step<Ctx>[] = [
{ name: "reserve-stock", run: reserveStock, compensate: releaseStock },
{ name: "authorise-payment", run: authorisePayment, compensate: voidAuthorisation },
{ name: "confirm-order", run: confirmOrder, compensate: cancelOrder },
];
async function runSaga(steps: Step<Ctx>[], ctx: Ctx) {
const done: Step<Ctx>[] = [];
for (const step of steps) {
try {
await step.run(ctx);
done.push(step); // persist this list: a crash mid-saga must be resumable
} catch (err) {
// Undo in reverse order. Retry each compensation until it succeeds, and make it idempotent,
// because the outage that failed the step may still be going on.
for (const prev of done.reverse()) await retryForever(() => prev.compensate(ctx));
throw err;
}
}
}
Worked failure: three steps, reserve stock, authorise payment, confirm order. If confirm order throws, the saga runs two compensations in reverse, void the authorisation and then release the stock. The worst case is 3 + 2 = 5 local transactions, and none of them holds a lock across a network call. Order the steps so the one that cannot be undone, such as sending a receipt, runs last. That is a rule of thumb of mine, not a theorem.
What does a saga give up? Semantic isolation problems
A saga is atomic only eventually, and it has no isolation. The Sagas paper describes sequences of transactions that can be interleaved with other work, so other transactions see the intermediate state. Between step 1 and a compensation, the reserved stock is visible to everyone.
Worked example: stock is 10 units. Saga A reserves 6, leaving 4. Saga B asks for 5, sees 4 and is told the item is out of stock. Then A's payment fails and its compensation returns the 6 units, 4 + 6 = 10, which would have covered B. B was rejected against a state that never became real. Three common countermeasures reduce the damage.
Semantic lock: mark the row with a status such as PENDING so other flows know the value is provisional and can wait or retry.
Commutative updates: prefer increment and decrement over writing an absolute value, so a compensation can be applied in any order without overwriting someone else's change.
Reread before write: check the current value or a version number immediately before updating it, instead of trusting what the saga read earlier.
Two-phase commit vs saga: which should you use?
Neither is better in general. They sit at different points on a trade between isolation and availability, and the comparison below is where I would start.
Dimension
Two-phase commit
Saga
Atomicity
All participants commit or none do, decided by one coordinator
Each step commits alone; the whole is atomic only eventually, through compensation
Isolation
Locks are held until the decision, so partial state stays hidden
None: partial state is visible, so you add semantic locks
When something fails
A coordinator crash leaves in-doubt participants blocked, holding locks
A failed step triggers compensations; no lock is held across steps
Participant requirements
Every participant must support prepare, such as PostgreSQL or an XA resource
Any service with a local transaction and an undo, including HTTP APIs and gateways
Best fit
A few databases you operate yourself, with short transactions
Cross-service or third-party flows and long running business processes
Main cost
Availability, plus prepared transactions someone must babysit
Design effort: compensations, idempotency and anomaly handling
Run through these questions in order and stop at the first decisive answer.
Can every participant do prepare? A payment gateway reached over HTTP cannot, which settles it in favour of a saga.
Could one database hold all the data? On a single VPS, one Postgres transaction beats both patterns, so ask this before reaching for either.
Can you write an undo for every step? If a step has no sensible compensation, move the service boundary instead of forcing a saga.
Can the business tolerate other users briefly seeing provisional state? If not, 2PC or a single database is the honest answer.
Who will watch the in-doubt work? Either pg_prepared_xacts for 2PC or the saga log for a saga needs an owner and an alert.
Start by trying to avoid the distributed transaction entirely. If it is unavoidable, choose two-phase commit only when you control every participant and can accept blocking, and choose a saga when steps cross services or vendors and you can write a safe undo for each one. Either way, nothing is finished until someone is alerted when a prepared transaction or a half-run saga is left behind.