Short answers to what readers ask most about this topic.
01What is the difference between synchronous and asynchronous communication in microservices?
In synchronous communication one service calls another's API over HTTP or gRPC and waits for the response. In asynchronous communication a service sends a message and carries on, and other services process it later through a broker. The first couples the two services in time; the second lets the receiver be down when the message is sent.
02When should I use asynchronous messaging instead of REST between services?
Use messaging when the caller does not need the result to continue: side effects such as receipts and notifications, fan-out to several consumers, long-running jobs, and bursts that a queue can level out. Keep REST or gRPC for queries where the user is waiting for an answer. Decide for each interaction rather than once for the whole system.
03How much does a chain of synchronous calls reduce availability?
If every service in the chain is required, availability is the product of the individual figures. Five services at 99.9 percent each give 0.999 to the fifth power, about 99.501 percent, or roughly 215.6 minutes of downtime in a 30-day month. The calculation assumes independent failures, so it is a guide to the shape rather than a forecast.
04Is HTTP always synchronous?
HTTP is a synchronous protocol because the sender waits for a response, even when the client uses non-blocking I/O. You can still build asynchronous behaviour on top of it: return 202 Accepted immediately and let the client poll a status endpoint or receive a webhook when the work finishes.
05What does 202 Accepted mean and when should I return it?
RFC 9110 defines 202 as accepted for processing but not completed, and calls the response intentionally noncommittal. Return it for work that takes longer than a request should, such as reports or exports, with a status URL the client can poll. Because the request might still fail later, the status endpoint must be able to report failure.
Sync vs Async Microservices Communication: How to Choose
When a service should call another and wait, and when it should publish a message instead: availability arithmetic, latency, failure handling, consistency, and the 202 Accepted hybrid.
Use synchronous calls when the caller needs the answer to continue, and asynchronous messaging for side effects, fan-out and slow work. Chained synchronous calls multiply availability: five services at 99.9 percent each give about 99.5 percent. Choose per interaction, and join the two with a 202 Accepted response plus a status endpoint.
The question comes up the first time a second service appears. A carwash POS finishes a payment, and something else has to happen: the receipt goes out, stock is adjusted, a report counter moves. Do you call each of those services and wait, or do you announce that the payment happened and let them react?
This post is the decision guide for that choice. I ran no benchmark for it. The definitions and trade-offs come from the Azure Architecture Center and microservices.io, the 202 behaviour from RFC 9110, and every figure below is arithmetic I derived and show in full, using assumed numbers that are labelled as assumptions.
What is the difference between synchronous and asynchronous communication?
In synchronous communication a service calls an API that another service exposes, over HTTP or gRPC, and the caller waits for the response. In asynchronous messaging a service sends a message without waiting for a reply, and one or more services process it later. The Azure Architecture Center draws one distinction worth keeping: asynchronous I/O is not an asynchronous protocol. An HTTP client can use non-blocking I/O and HTTP is still a synchronous protocol, because the sender waits for an answer.
Dimension
Synchronous request/response
Asynchronous messaging
Does the caller wait?
Yes, for the response
No, it continues after the hand-off
Coupling in time
Both services must be up at the same moment
The consumer can be down; the broker holds the message
If the callee is down
The caller gets an error or a timeout
The sender still succeeds; messages are processed on recovery
Latency the user sees
The sum of every hop in the chain
Only the hand-off; the work finishes later
Consistency
Immediate answer, simple to reason about
Eventual; consumers must tolerate duplicates
Extra moving part
None, apart from the network
A broker that itself must stay available
The table is the whole argument in compressed form. Synchronous calls buy you an immediate, simple answer and cost you coupling in time. Messaging buys you independence and costs you a broker, duplicate handling and a harder way to get a response back. The Azure guide lists the same costs: complexity, queue latency when queues fill, and the difficulty of building request-response on top of messaging.
How much availability does a synchronous call chain cost?
If a request needs every service in a chain to answer, the availability of the chain is the product of the availabilities. Microservices.io states the same drawback in words: with remote procedure invocation, the client and the service must both be available for the duration of the interaction. The arithmetic assumes failures are independent and every hop is mandatory, so treat the numbers as an upper bound for the shape, not a forecast for your system.
Services in the chain, each at 99.9 percent
Chain availability
Downtime per 30 days
1
99.900 percent
43.2 minutes
2
99.800 percent
86.4 minutes
3
99.700 percent
129.5 minutes
5
99.501 percent
215.6 minutes
10
99.004 percent
430.1 minutes
20
98.019 percent
855.8 minutes
The code computes the table. A 30-day month has 43,200 minutes, and downtime is the unavailable share of that.
// availability.ts - run with: node availability.ts (Node 22.18+ strips the types itself)
const MINUTES_PER_30_DAYS = 30 * 24 * 60; // 43,200
function chain(perService: number, n: number) {
// Every hop is mandatory and failures are assumed independent, so availabilities multiply.
const availability = perService ** n;
return {
n,
availabilityPct: (availability * 100).toFixed(3),
downtimeMinutes: ((1 - availability) * MINUTES_PER_30_DAYS).toFixed(1),
};
}
for (const n of [1, 2, 3, 5, 10, 20]) console.log(chain(0.999, n));
// n=1 99.900 43.2
// n=2 99.800 86.4
// n=3 99.700 129.5
// n=5 99.501 215.6
// n=10 99.004 430.1
// n=20 98.019 855.8
Five services that each look healthy at 99.9 percent give a user path with roughly five times the downtime of one of them. Messaging changes the shape of this sum. The user-facing request depends on the services it must have right now plus the broker, while the rest are fed from the queue and can be down without failing the request. The broker is now a dependency with its own availability, which microservices.io lists as the main drawback of messaging, so on a single VPS it only helps if it is monitored as seriously as the database.
How does latency add up in a synchronous chain?
Latency in a sequential chain is the sum of the hops, and the slowest dependency sets the floor for the whole request. The Azure guide calls out the same effect: if service A calls B, which calls C, waiting on synchronous calls can add unacceptable latency. Timeouts and retries then make the worst case far worse than the typical one.
Every figure in this snippet is an assumption chosen to show the arithmetic, not a measurement.
// Assumed: five sequential hops, 40 ms each. Latency adds: 5 * 40 = 200 ms
// before the caller has done any work of its own.
// Assumed: a 2000 ms timeout and 3 attempts per call.
// Worst case at ONE layer = 3 * 2000 = 6000 ms, plus the backoff waits between attempts.
// Assumed: 3 layers, each allowing 3 attempts against the layer below.
// Calls reaching the deepest service for one user action = 3 * 3 * 3 = 27.
// Right: one deadline for the whole request, each hop spends only what is left.
const DEADLINE_MS = 1500;
const startedAt = Date.now();
const remaining = () => Math.max(0, DEADLINE_MS - (Date.now() - startedAt));
// pass remaining() as the timeout of the NEXT call, never a fresh fixed number
The practical rule is to give the top-level request one deadline and to let each hop spend only what is left of it. Retry policy deserves its own design, which the post on retries with exponential backoff and jitter covers, and the bulkhead post covers how to stop one slow dependency from holding every connection.
Retries multiply across layers. If three layers each allow three attempts, the deepest service can receive 27 calls for one user action, and it is already the one struggling. Retry at one layer, and only for idempotent operations.
How do failures differ: timeouts and retries or a durable queue?
A synchronous caller owns the failure. It must set a timeout, decide what to retry, and decide what the user sees when the callee does not answer. The first snippet is a NestJS client call with an explicit timeout. Without one you are relying on a default, and a hung pricing service then hangs the checkout that called it.
import { HttpService } from "@nestjs/axios";
import { Injectable, ServiceUnavailableException } from "@nestjs/common";
import { firstValueFrom } from "rxjs";
@Injectable()
export class PricingClient {
constructor(private readonly http: HttpService) {}
async quote(serviceId: string): Promise<{ priceIdr: number }> {
try {
// Right: an explicit timeout. Wrong: relying on a default and waiting on a hung callee.
const { data } = await firstValueFrom(
this.http.get("http://pricing:3000/quotes/" + serviceId, { timeout: 800 }),
);
return data;
} catch {
// The caller owns this failure: decide what the cashier sees.
throw new ServiceUnavailableException("Pricing is unavailable, try again");
}
}
}
A publisher hands the problem to the broker. If the consumer is down, the Azure guide notes that the sender can still send, and the messages are picked up when the consumer recovers. The second snippet publishes an event instead. The caller no longer knows or cares who reacts.
import { Inject, Injectable } from "@nestjs/common";
import { ClientProxy } from "@nestjs/microservices";
import { lastValueFrom } from "rxjs";
@Injectable()
export class PaymentEvents {
constructor(@Inject("EVENT_BUS") private readonly bus: ClientProxy) {}
async paymentCaptured(paymentId: string, amountIdr: number) {
// The receipt, stock and report services react on their own schedule.
// Include an id: delivery is at-least-once, so consumers must de-duplicate on it.
await lastValueFrom(
this.bus.emit("payment.captured", { eventId: paymentId, amountIdr }),
);
}
}
The price is that delivery usually becomes at-least-once. The Azure guide says you must handle duplicated messages either by de-duplicating or by making operations idempotent. The queue is a durable buffer, not a guarantee that work happens exactly once.
What does asynchronous messaging do to consistency?
A synchronous call returns when the work is done, so the next line of code can rely on it. After a publish, the work has only been handed over. Other services will be briefly out of date, and the UI has to be honest about that, for example by showing a pending state instead of a confirmed one. That is eventual consistency, and it is a product decision as much as a technical one.
The classic trap is the dual write: commit to the database, then publish, and crash between the two. The state changed and nobody was told. Writing the event to an outbox table in the same transaction as the state change, and publishing from that table, closes the gap.
// Wrong: if the process dies after the commit, the payment exists and nobody is told.
await db.transaction((tx) => tx.insert(payments).values(row));
await bus.emit("payment.captured", row);
// Right: the event is written in the SAME transaction; a relay publishes it afterwards.
await db.transaction(async (tx) => {
await tx.insert(payments).values(row);
await tx.insert(outbox).values({
topic: "payment.captured",
payload: JSON.stringify(row),
publishedAt: null, // the relay sets this after a successful publish
});
});
The outbox makes the publish at-least-once, which is why the consumer-side de-duplication from the previous section is not optional.
When is synchronous right, and when is asynchronous right?
Decide per interaction, not per system. Ask one question of each call: does the caller need the result to continue? The split I would draw in a carwash POS follows from it.
Synchronous fits when the answer is needed now:
Reading a price or a customer record while building the checkout screen.
Validating a voucher before the total is shown.
Any query where an error should be shown to the user immediately.
Operations that must succeed or fail together before you reply.
Asynchronous fits when the work can happen after the reply:
Side effects such as sending a receipt or a notification.
Fan-out, where one payment feeds stock, reporting and analytics.
Long-running work such as a monthly report or an export.
Load levelling, where a queue absorbs a burst that a downstream service drains at its own pace.
If you are still choosing a protocol for the synchronous half, the earlier gRPC versus REST comparison covers it. For the asynchronous half, the posts on message queues and on pub/sub versus queues cover when a competing-consumer queue or a fan-out topic fits. And if events feel like too much machinery, the post on when not to go event-driven is the honest counterweight.
Can you combine both with a 202 Accepted response?
Yes, and it is the most useful hybrid. The client sends a synchronous request, the server validates it, stores the job and returns 202 Accepted at once. RFC 9110 defines it as accepted for processing, but not completed, and describes the response as intentionally noncommittal. It also says the response should describe the request's current status and point to a status monitor, which is the polling endpoint below.
import { Body, Controller, Get, Headers, HttpCode, NotFoundException, Param, Post, Res } from "@nestjs/common";
import type { Response } from "express";
@Controller("reports")
export class ReportsController {
constructor(private readonly reports: ReportsService) {}
@Post()
@HttpCode(202) // accepted for processing, not completed (RFC 9110, 15.3.3)
async create(
@Body() dto: CreateReportDto,
@Headers("idempotency-key") key: string,
@Res({ passthrough: true }) res: Response,
) {
// Insert the job row and enqueue in ONE transaction; same key returns the same job.
const job = await this.reports.enqueue(dto, key);
res.setHeader("Location", "/reports/" + job.id); // the status monitor
res.setHeader("Retry-After", "5"); // suggested seconds before polling
return { id: job.id, status: "queued" };
}
@Get(":id")
async status(@Param("id") id: string) {
const job = await this.reports.find(id);
if (!job) throw new NotFoundException();
// queued | running | done | failed - a 202 promised nothing, so failure must be reportable.
return { id: job.id, status: job.status, error: job.error ?? null, resultUrl: job.resultUrl ?? null };
}
}
The client polls the status URL, or you call a webhook when the job finishes. The create and the job row are written together, so the 202 is never a promise the server cannot keep. Add an idempotency key on the POST so a retried request does not queue the work twice.
RFC 9110 is explicit about the limit: there is no facility in HTTP for re-sending a status code from an asynchronous operation. A 202 says the request was taken, not that it will succeed, so the status endpoint must be able to report failure.
Tag each interaction in a design doc as query, command needing a result, or side effect. Queries stay synchronous, side effects become events, and commands with slow results get the 202 shape. The system ends up mixed, which is the correct outcome.
Do not pick a style for the whole architecture. Keep synchronous calls for the answers a caller needs right now, and count the chain: every mandatory hop multiplies availability and adds latency. Move side effects, fan-out and slow work behind a durable message, accept at-least-once delivery, and return 202 with a status endpoint where a client still wants to know the outcome.