Short answers to what readers ask most about this topic.
01What is the difference between a service mesh and an API gateway?
An API gateway handles north-south traffic: external clients calling into your system, where it does authentication, rate limiting and routing. A service mesh handles east-west traffic between your own services, adding mutual TLS, retries and telemetry through proxies. They overlap in routing but serve different trust boundaries.
02Do I need a service mesh for microservices?
Not by default. With a small number of services owned by one team, application libraries for timeouts, retries and logging cost less than operating a control plane and proxies. A mesh starts to pay off when several teams need uniform encryption, tracing and traffic control without editing each codebase.
03What is the difference between sidecar and ambient mesh?
A sidecar mesh runs a proxy in every application pod. Istio's ambient mode instead uses a per-node ztunnel for mutual TLS and L4 features, plus optional per-namespace waypoint proxies for L7 features. The Istio announcement for 1.24 marked its core ambient components Stable.
04Can I use an API gateway and a service mesh together?
Yes, and many larger systems do. The gateway validates external callers and forwards requests inward, then the mesh encrypts and manages each internal hop. Decide which layer owns retries so they do not multiply across layers.
05Is nginx an API gateway or a reverse proxy?
Nginx is a reverse proxy that can serve as a simple API gateway. With proxy_pass, upstream groups and limit_req it covers routing and rate limiting for a small system. Features such as developer portals or key management usually need a dedicated gateway product or application code.
Service Mesh vs API Gateway: Differences and When to Use Each
Service mesh vs API gateway: north-south versus east-west traffic, sidecar versus ambient mesh, real nginx and Istio config, cost, and when you need neither.
An API gateway handles north-south traffic: it sits at the edge, authenticates clients, rate limits and routes requests into your system. A service mesh handles east-west traffic: it adds mutual TLS, retries and telemetry between services through proxies. A small team on one server usually needs neither, only a reverse proxy.
Every few months someone asks whether a system needs Istio, and the honest first question is what problem they are trying to solve. The two tools get confused because both are proxies that route HTTP, both can enforce policy, and both show up in the same architecture diagrams.
This post separates them by the direction of the traffic they handle, shows a real nginx gateway config and real Istio YAML, checks the sidecar and ambient models against the Istio and Linkerd documentation, and ends with a decision checklist. My own systems are a Qilap-style ERP and POS on NestJS, Postgres and Redis in Docker on a single VPS, which is exactly the situation where the answer is usually neither.
What is the difference between north-south and east-west traffic?
North-south traffic crosses the boundary of your system: a browser, a mobile app or a partner calls in from outside. East-west traffic stays inside: the orders service calls the catalog service, a worker calls the billing API. The terms are network-diagram slang rather than a formal standard, but they map cleanly onto the two tools.
An API gateway owns north-south. It is the single front door where you terminate TLS, identify the caller, throttle abusive clients and decide which backend a path belongs to. A service mesh owns east-west. It assumes the caller is already inside and focuses on making every internal hop encrypted, observable and resilient without changing application code.
What does an API gateway do at the edge?
Gateway work is about outsiders: authentication of tokens or API keys, rate limiting, request routing by path or host, TLS termination, and sometimes response caching or request transformation. In its simplest form that is a reverse proxy. The nginx config below routes two path prefixes to two upstream groups and rate limits each client address, using the directives from the official limit_req documentation.
# /etc/nginx/conf.d/gateway.conf (north-south: the one door into the system)
# 10 requests/second per client IP on average, bursts of up to 20 extra.
limit_req_zone $binary_remote_addr zone=api:10m rate=10r/s;
limit_req_status 429; # default is 503, which reads like an outage
upstream orders_api { server 10.0.0.11:3001; server 10.0.0.12:3001; }
upstream catalog_api { server 10.0.0.21:3002; }
server {
listen 443 ssl;
server_name api.example.com;
location /orders/ {
limit_req zone=api burst=20 nodelay; # nodelay: serve the burst now
proxy_set_header X-Request-Id $request_id;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_pass http://orders_api/;
}
location /catalog/ {
limit_req zone=api burst=20 nodelay;
proxy_pass http://catalog_api/;
}
}
The arithmetic is worth doing once. A rate of 10r/s with burst=20 and nodelay means a client averages 10 requests per second but can fire 20 extra immediately before being rejected with the 429 status; the nginx documentation notes the default rejection code is 503, which is why the config overrides it. Authentication usually sits in front of this block, either as a JWT check in the gateway or as an nginx subrequest to an auth service. The gateway pattern article linked at the end shows the NestJS version.
What does a service mesh do between services?
A mesh moves cross-cutting concerns for internal calls out of application code and into proxies: mutual TLS between workloads, retries and timeouts, traffic splitting for canary releases, and uniform metrics and traces. Istio expresses these as Kubernetes resources. The first policy refuses plaintext traffic inside one namespace, and the second adds a timeout, retries and a 90/10 canary split for the orders service.
# East-west: refuse any plaintext call between workloads in "shop".
apiVersion: security.istio.io/v1
kind: PeerAuthentication
metadata:
name: default # "default" = namespace-wide policy
namespace: shop
spec:
mtls:
mode: STRICT # PERMISSIVE accepts both; use it while migrating
Field names above follow the Istio reference pages for PeerAuthentication and VirtualService, where the documented default for retries is 2 attempts. Linkerd takes a narrower route: its documentation says it enables mutual TLS automatically for TCP traffic between meshed pods, with certificates bound to the pod ServiceAccount identity and expiring after 24 hours. Both approaches get you encrypted, identity-checked internal calls without touching service code.
Retries multiply. If your HTTP client retries 3 times and the mesh also retries 3 times, one failing user request can hit the broken service 9 times, and a gateway retry on top makes it 27. Pick one layer to own retries, and pair it with timeouts and backpressure so a struggling service is not buried by its own callers.
What are the sidecar and ambient mesh models?
The classic mesh injects a proxy next to every application pod. Linkerd documents this as a data plane proxy in each pod, and Istio's sidecar mode uses an Envoy proxy the same way. Istio's ambient mode removes that per-pod proxy and splits the data plane into two layers, according to the Istio ambient overview:
ztunnel: a per-node proxy written in Rust that provides mutual TLS, authentication, L4 authorisation and L4 telemetry. It does not parse HTTP headers.
Waypoint proxy: an optional Envoy deployment, per namespace, installed and scaled separately from your apps. L7 features such as VirtualService routing and L7 authorisation need it.
Coexistence: sidecar and ambient workloads can live in the same mesh, so adoption can be incremental.
Ambient is version-dependent: the Istio announcement for release 1.24 says its core components were marked Stable then. The overview page gives no overhead numbers, only that waypoints are heavier than ztunnel alone, so treat any claimed percentage saving as something to measure on your own cluster rather than a fact.
Service mesh vs API gateway: how do they compare?
The table summarises where each tool sits and what it is for. Read the rows as a division of labour rather than a ranking, because most production systems that use both are using each for the job it is built for.
Aspect
API gateway
Service mesh
Traffic direction
North-south, from clients into the system
East-west, between internal services
Where it runs
One or a few edge proxies
A proxy per pod (sidecar) or per node plus optional waypoints (ambient)
Authentication
Identifies external callers: tokens, API keys
Identifies workloads: mutual TLS identities
Rate limiting
Core feature, per client or key
Possible, but rarely the main reason to adopt a mesh
Resilience
Timeouts and upstream failover at the edge
Retries, timeouts and traffic splitting on every internal hop
Observability
Edge request logs and metrics
Uniform metrics and traces for every service-to-service call
Operational cost
Low: one config file can be enough
High: a control plane, proxy upgrades and new failure modes
The overlap is real. A gateway can retry upstream calls and a mesh can rate limit, so the useful question is which layer should own a concern. Put anything about untrusted callers at the gateway, and anything about trusted internal hops in the mesh or, for a small system, in the application itself.
How much does a service mesh cost in complexity?
The cost is a count of moving parts you now operate. Take a hypothetical cluster of 12 services with 3 replicas each on 4 nodes in 2 namespaces. The sketch below shows how many proxies each model implies, and how stacked retries inflate load. The numbers are arithmetic on assumed inputs, not a benchmark.
// Hypothetical cluster: 12 services x 3 replicas on 4 nodes, 2 namespaces
pods = 12 * 3 // 36
sidecar proxies = pods // 36 (one Envoy per pod)
ambient L4 only = nodes // 4 (one ztunnel per node)
ambient + L7 = 4 + 2 // 6 (ztunnel per node + one waypoint per namespace)
// Retry stacking: client library x mesh
worst case calls per user request = 3 * 3 // 9 hits on the failing service
Sidecar mode gives 36 proxies to upgrade and restart alongside the apps, while ambient with waypoints gives 6, though those 6 are shared infrastructure with their own scaling. Either way you also run a control plane, learn a new set of resources, and debug failures where the proxy rather than your code is at fault. Service discovery, which another article in this series covers, is something you still need regardless of the mesh.
Adopt in stages. Start with PERMISSIVE mutual TLS, watch the traffic, then switch to STRICT per namespace. In ambient mode you can begin with the L4 ztunnel only and add a waypoint just for the namespaces that need L7 routing.
Does a small team need a service mesh or an API gateway?
Usually not a mesh, and often not even a full gateway. A single VPS running a NestJS API, Postgres, Redis and a few workers in Docker has one real edge, and a reverse proxy handles it. Use this checklist before adding either tool:
Do you have fewer than roughly a dozen services owned by one team? Then app-level libraries for timeouts, retries and logging are cheaper than a mesh.
Does anything traverse an untrusted network between services? If all traffic stays on one host or a private network, mutual TLS is a smaller win.
Do you need per-client rate limits, API keys or request routing for external callers? That is a gateway job, and nginx or a NestJS gateway covers it.
Do multiple teams need uniform encryption, tracing and canary releases without changing each codebase? That is where a mesh starts to pay for itself.
If you answered yes only to the third question, you want a gateway. If the first answer is yes and the rest are no, you want a reverse proxy and good libraries, nothing more.
Can a service mesh and an API gateway work together?
Yes, and this is the common arrangement in larger systems. The gateway receives the external request, validates the caller and forwards it with a request ID. From that point the mesh takes over, encrypting each internal hop and applying the retry and routing policy. The order to follow when adopting them is:
Start with a reverse proxy at the edge for TLS, routing and rate limits.
Add the application-level basics: timeouts, bounded retries and structured logs with a request ID.
Introduce a mesh only when the checklist shows several teams repeating the same internal-traffic work, and begin with mutual TLS in permissive mode.
Kubernetes users can also model the edge with the Gateway API, which is covered in a separate comparison with Ingress. Some meshes can act as the edge gateway too, but keeping the two roles named separately keeps the policy easier to reason about.
The rule I carry is to choose by traffic direction. Untrusted callers entering the system need a gateway. Trusted services talking to each other need a mesh only once the repeated work across teams costs more than operating the proxies. Until then, a reverse proxy and disciplined libraries are the right size.