Short answers to what readers ask most about this topic.
01How does service discovery work in microservices?
Each instance registers its address with a registry, renews that registration with a heartbeat, and is removed when it deregisters or its TTL runs out. A caller then looks up the service by name at call time instead of using a fixed IP. The lookup is done either by the client itself or by a router in front of the instances.
02What is the difference between client-side and server-side service discovery?
In client-side discovery the calling service queries the registry and picks an instance itself, which saves a network hop but needs discovery logic in every language you use. In server-side discovery the client calls a router or load balancer that queries the registry for it, which keeps clients simple but adds a component and an extra hop.
03How do Docker Compose services find each other?
Compose attaches your services to a default network and registers each service name with an internal DNS server. A container can then call http://orders:8080 and the name resolves to the current container IP. Container IPs are assigned dynamically, so the name is the part you should rely on.
04How does service discovery work in Kubernetes?
A normal Service gets a DNS name of the form my-svc.my-namespace.svc.cluster-domain.example that resolves to the Service's cluster IP, and kubelet configures each Pod with a search list so short names work inside the same namespace. CoreDNS is the default implementation of cluster DNS. A headless Service skips the cluster IP and resolves to the IPs of all its Pods.
05Do I need Consul or etcd for service discovery?
Not on a single host, where Compose service names or a reverse proxy are enough. Consul and etcd earn their place when you span several machines, need health-aware routing that plain DNS cannot express, or want change notifications through leases and watches. Before then they are another stateful system to run.
How does service discovery work in microservices? Registry, heartbeat and TTL, client-side vs server-side, DNS in Docker Compose and Kubernetes, Consul.
Service discovery lets one microservice find another's current address at call time instead of hardcoding an IP. Instances register with a registry, renew a TTL heartbeat, and drop out when they stop or fail health checks. Callers resolve a name through DNS, as Docker Compose and Kubernetes do, or through a registry such as Consul or etcd.
The quickest way to connect two services is to paste one container's IP into the other's environment variable. It works until the container restarts and the address is no longer the same one. Nothing is wrong in the code. The address you wrote down has simply stopped being true.
This post answers one question: how does service discovery work in microservices? It covers why static addresses fail, how a registry handles register, heartbeat and deregister, client-side against server-side discovery, DNS in Docker Compose and Kubernetes, and where Consul and etcd fit. The Docker output and the TypeScript registry output are real runs I made for this article, on Docker 29.8.1 and Node 26. Everything else traces to the official docs listed at the end.
Why do static IPs and hardcoded hosts break microservices?
Because the address of a service is not a property of the service. It is a property of wherever the scheduler put it this time. The Docker Compose networking docs say container IPs are assigned dynamically when a container starts and are not kept, and that the service name is the stable part. I reproduced the failure: I stopped the orders container, let a second container take its freed address, then started orders again.
$ docker compose up -d orders
before: 172.18.0.2
$ docker compose stop orders
$ docker run -d --network sdemo_default --name squatter busybox sleep 60
squatter took: 172.18.0.2 # something else grabbed the freed address
$ docker compose start orders
after: 172.18.0.3 # same service, new IP
$ docker run --rm --network sdemo_default busybox nslookup orders
Name: orders
Address: 172.18.0.3 # the NAME followed it, a hardcoded IP would not
One honest detail: a plain force-recreate with nothing else on the network gave orders the same address back in my run, so the drift is not guaranteed on every restart. It happens when anything else, such as another container, a scale-out or a deploy, takes the address in between. With one container on one host you may live with a hardcoded IP for months and then lose an afternoon to it, which is worse than failing immediately.
A hardcoded IP fails silently and late. The service is healthy, the config looks right, and requests go to a different container or to nothing. Treat any IP literal in service config as a bug waiting for a restart.
What is a service registry, and how do register, heartbeat and deregister work?
A service registry is a table of service name to live instances. Each instance registers itself on startup with an id, host and port, renews that registration on a timer, and deregisters on graceful shutdown. If it crashes it cannot deregister, so the registration carries a TTL and the registry expires it on its own. This is the whole mechanism behind Consul TTL checks and etcd leases, reduced to its smallest TypeScript form.
type Instance = { id: string; host: string; port: number; expiresAt: number };
const TTL_MS = 10_000;
export class Registry {
private byService = new Map<string, Map<string, Instance>>();
// register and heartbeat are the same call: upsert, push expiry forward.
register(service: string, id: string, host: string, port: number, now = Date.now()) {
const instances = this.byService.get(service) ?? new Map<string, Instance>();
instances.set(id, { id, host, port, expiresAt: now + TTL_MS });
this.byService.set(service, instances);
}
// Graceful shutdown path: remove immediately instead of waiting out the TTL.
deregister(service: string, id: string) {
this.byService.get(service)?.delete(id);
}
// Lazy expiry: an entry is dead the moment now passes expiresAt,
// whether or not a sweeper has deleted it yet.
lookup(service: string, now = Date.now()): Instance[] {
const instances = this.byService.get(service);
if (!instances) return [];
for (const [id, inst] of instances) {
if (inst.expiresAt <= now) instances.delete(id);
}
return [...instances.values()];
}
}
I ran this with a fake clock so the timeline is exact. Both instances register at t=0 with a 10 second TTL. Only orders-a sends a heartbeat, at 6 seconds.
// fake clock, so the run is instant: orders-a and orders-b register at t=0
// orders-a heartbeats at t=6s, orders-b goes silent
t=6s [ 'orders-a', 'orders-b' ]
t=10s [ 'orders-a' ] // orders-b: 0 + 10 s TTL, gone exactly now
t=16s [] // orders-a: 6 s + 10 s TTL
The arithmetic is the point. orders-b was last seen at 0, so it expires at 0 plus 10 equals 10 seconds. orders-a renewed at 6, so it lives until 6 plus 10 equals 16 seconds. Note that register and heartbeat are one call here: an upsert that pushes expiry forward. That is also how a recovering instance rejoins without any special path.
Heartbeat at about one third of the TTL. With a 10 second TTL that is every 3 seconds, so beats land at 3, 6 and 9 seconds before expiry and one network blip, even two, does not evict a healthy instance. This ratio is my own rule of thumb, not a number from any vendor doc.
Client-side or server-side discovery: which should I use?
The two patterns differ in who queries the registry. In client-side discovery, the client asks the registry for instances and picks one itself. In server-side discovery, the client calls a router or load balancer that queries the registry and forwards the request. The microservices.io pattern pages list the trade-offs, and the table maps them.
Aspect
Client-side discovery
Server-side discovery
Who queries the registry
The calling service, in its own code
A router or load balancer in front of the instances
Network hops
Fewer, the client connects directly
More, every request passes through the router
Client code
Needs discovery logic per language and framework
Simpler, the client just calls one address
Extra component to run
None beyond the registry itself
The router must be installed, configured and kept up
Real-world example
An app calling the Consul or etcd API directly
AWS Elastic Load Balancer, or the Kubernetes proxy on each host
For a small estate I lean to server-side: the logic lives in one place and every service stays dumb about the network. The pieces already exist in most stacks, such as a reverse proxy or the cluster itself. I have written separately about the API gateway pattern in NestJS and a production load balancer config, which are the usual server-side routers. A service mesh moves the client-side logic into a sidecar so you keep the benefits without per-language code, which is its own comparison.
How does DNS-based discovery work in Docker Compose and Kubernetes?
DNS is the lowest-friction discovery there is, because every runtime already knows how to resolve a name. The Compose docs say each service registers its name with an internal DNS server on the app network, so a container can look up web or db and get the right container's IP. Here is a two-service file where pos calls orders by name. No ports are published, and no IP appears anywhere.
services:
orders:
image: busybox:latest
# One static file on 8080. No published ports: only the Compose network reaches it.
command:
- sh
- -c
- mkdir -p /www && echo 'orders says wash 8841 is done' > /www/index.html && httpd -f -p 8080 -h /www
pos:
image: busybox:latest
depends_on: [orders]
# Resolve and call the other service by NAME, never by IP.
command:
- sh
- -c
- sleep 2; nslookup orders; wget -qO- http://orders:8080/
The output below is pasted from a real docker compose up. The resolver is 127.0.0.11, Docker's embedded DNS, and the name orders became 172.18.0.2 on the project network.
One caveat from the Docker bridge docs: automatic name resolution works on user-defined networks, which Compose creates for you, but on the default bridge network containers cannot resolve each other by name. Kubernetes does the same job at cluster scale. A normal Service gets an A or AAAA record of the form my-svc.my-namespace.svc.cluster-domain.example that resolves to the Service's cluster IP, and kubelet writes a search list into every Pod so short names work.
# /etc/resolv.conf inside a Pod (the example from the Kubernetes DNS page)
nameserver 10.32.0.10
search <namespace>.svc.cluster.local svc.cluster.local cluster.local
options ndots:5
# So a Pod in namespace "test" can call:
# http://orders (same namespace, short name)
# http://orders.prod (other namespace)
# http://orders.prod.svc.cluster.local (fully qualified)
CoreDNS is the default implementation of Kubernetes cluster DNS, and its kubernetes plugin watches the cluster API for endpoints. Its answers carry a TTL that defaults to 5 seconds, with a maximum of 3600 and a value of 0 meaning no caching. A headless Service, one with no cluster IP, resolves to the IPs of all its Pods instead of one virtual address, and clients are expected to pick from that set themselves.
When do I need Consul or etcd instead of DNS?
Plain DNS answers where a service is, but it says little about whether an instance is healthy right now, and it gives you no way to be told when the set changes. etcd gives you the primitives to build that. Its leases have a TTL, keys attached to a lease are deleted when the lease expires or is revoked, and a client must send keepalive requests to hold the lease. Each expired key generates a delete event, and the Watch API streams those events, so a client can react to a change instead of polling for it.
Consul packages the same idea as a product. Its service catalog is the source of truth for where services are, and health checks remove unhealthy instances from the catalog so they stop receiving requests. Per the Consul docs, a TTL check waits for an external process to report to the agent check endpoint, and if no update arrives before the ttl duration the service is marked critical. The definition is short.
service {
name = "orders"
id = "orders-a"
port = 3000
check {
id = "orders-a-ttl"
name = "heartbeat"
ttl = "10s" # no update within 10 s and the check turns critical
}
}
My position: do not add Consul or etcd to a single host. They earn their place when you run across several machines, need health-aware routing that DNS alone cannot express, or need change notifications. Before that point they are one more stateful system to keep alive, and a registry that is down takes discovery down with it.
What goes wrong with health checks, stale entries and caching?
A registry can only be as fresh as its slowest layer. Four failure shapes cover most incidents.
Stale entry: a crashed instance stays listed until its TTL runs out, and callers keep sending it traffic in the meantime.
Shallow health check: the process answers but its database connection is gone, so the instance reports healthy and fails every real request.
Client-side caching: the caller keeps using a resolved address for the full cache lifetime after the registry has already removed it.
Flapping: a TTL that is too tight against the heartbeat interval evicts healthy instances during a brief pause, then re-adds them, which churns every client.
The delays stack, so add them. If the registry TTL is 10 seconds and a client caches a resolved address for 5 seconds, matching the CoreDNS default, a crashed instance can still receive traffic for 10 plus 5 equals 15 seconds in the worst case. Shortening the TTL narrows that window, but a tighter TTL needs a faster heartbeat, which means more registry load.
registry TTL = 10 s (entry survives 10 s after the last heartbeat)
client-side cache lifetime = 5 s (the CoreDNS kubernetes plugin default TTL)
worst case for traffic to a crashed instance
= TTL + cache = 10 + 5 = 15 s
heartbeat every TTL / 3 = 10 / 3 = 3.33 s -> use 3 s
beats at 3, 6, 9 s before the 10 s expiry -> two lost beats are survivable
Do not paper over stale entries by retrying forever. A retry against a dead address burns the caller's timeout budget each time. Pair discovery with a short connect timeout, and retry against a different instance rather than the same one.
What should I pick for a single VPS?
I run Docker on a single VPS, and that shapes this advice. Use this checklist in order and stop at the first step that covers you.
Put every service on one user-defined network, which Compose does by default, and call by service name. Do not publish ports you do not need.
Check that no IP literal appears in your environment variables or config files. Search for them before you deploy.
Add a real health check that exercises the dependencies, then let your orchestrator or reverse proxy stop routing to failing instances.
Move to Kubernetes Services and CoreDNS only when you outgrow one host, and keep the same name-based calling habit.
Add Consul or etcd only when you can name the requirement DNS cannot meet, such as health-aware routing across machines or change notifications.
A service's address is data that changes, so look it up at call time and call by name. Everything else in this post is a decision about how fresh that lookup must be and who performs it. On one host, Compose service names are enough. Past that, the registry, its TTL and every cache in front of it add up to how long you route to dead instances, so add them up before you tune anything.