Short answers to what readers ask most about this topic.
01How do you migrate a monolith to microservices with the strangler fig pattern?
Put a facade or reverse proxy in front of the monolith, then extract one slice at a time into a new service and route that path to it. Move the data the slice owns with an outbox or change data capture, ramp traffic by percentage, and keep a flag so you can route back. Repeat until the remaining monolith is small and stable.
02What is the difference between the strangler fig pattern and a big-bang rewrite?
A big-bang rewrite delivers all its value and all its risk on one cut-over day. The strangler pattern delivers one reversible slice at a time while the old system keeps running. A mistake costs one slice instead of the whole system.
03How do you handle the database when splitting a monolith?
Make exactly one system own each piece of data. Avoid dual writes, because the two writes are not in one transaction and can diverge silently. Use a transactional outbox or change data capture so the new service builds its own copy from events, then move write ownership once the copies reconcile.
04Can I do a gradual percentage rollout with nginx?
Yes. The split_clients directive hashes a string, such as the client address, and maps ranges of the hash to the percentages you list. The same client keeps landing on the same backend, and rolling back is changing the percentage to 0 and reloading nginx.
05When should I not migrate to microservices?
Do not migrate if your team is small and the real problem is slow releases, if the code has no clear module boundaries, if you cannot observe the current system, or if the motivation is fashion rather than a measured problem. A modular monolith is often a cheaper fix. Microservices add network failures and distributed data that a single-server setup rarely repays.
Strangler Fig Pattern: Monolith to Microservices Migration
How to migrate a monolith to microservices with the strangler fig pattern: a routing facade, one slice at a time, data ownership, rollback and when to stop.
The strangler fig pattern migrates a monolith to microservices without a rewrite. Put a facade in front of the system, route one path at a time to a new service, give that service its own data via an outbox or change data capture, keep a routing flag for instant rollback, and stop when the remainder is stable.
Most ERP and POS monoliths are not bad code. They are working code that nobody dares to touch, because the invoice screen, the stock ledger and the receipt printer all share one codebase and one database. The tempting answer is a rewrite. The usual outcome of a rewrite is two systems to maintain for a long time and a cut-over weekend that everyone dreads.
This post walks through the alternative: the strangler fig pattern. It covers the routing facade, the order in which to extract slices, the hard part (who owns the data), branch by abstraction, the anti-corruption layer, rollback, and the signals that mean you should not migrate at all. The pattern comes from Martin Fowler and the Azure Architecture Center, and the nginx directives are checked against the nginx documentation. The examples are illustrative worked examples, not measurements from a production migration.
What is the strangler fig pattern, and why not just rewrite?
Martin Fowler named the pattern after strangler figs, vines that grow around a host tree and gradually replace it. Applied to software, you build the new system around the edges of the old one and let it grow until the old one can be switched off. The Azure Architecture Center describes the same idea as incrementally replacing specific pieces of functionality with new applications and services, while a facade routes requests to either the legacy or the new implementation.
The reason this beats a rewrite is risk shape. A rewrite delivers all of its value, and all of its risk, on one day. A strangler migration delivers a small slice at a time, each one reversible, so a mistake costs one slice rather than the whole system. If you are still deciding whether to split at all, read the comparison in the microservices versus monolith post first, because the pattern assumes you have already decided the destination is separate services.
How does the facade route requests per path?
The facade is a reverse proxy that every client talks to. Initially it forwards everything to the monolith, so nothing changes for users. When the first service is ready you add one path rule. The config below uses nginx, which fits a single VPS with Docker: a path rule decides what is eligible to move, and a split_clients block decides what share of that traffic actually goes to the new service.
# /etc/nginx/conf.d/facade.conf (the strangler facade)
upstream legacy_erp { server 127.0.0.1:3000; } # the monolith, untouched
upstream orders_svc { server 127.0.0.1:4001; } # the first extracted slice
# Percentage split. split_clients hashes the string below, so the same client
# lands in the same bucket on every request: no flip-flopping mid-session.
# Rollback = change 10% to 0% and reload. No deploy needed.
split_clients "$remote_addr-orders" $orders_backend {
10% orders_svc;
* legacy_erp;
}
server {
listen 80;
server_name erp.example.com;
# Path rule: only /api/orders/ is eligible to leave the monolith.
location /api/orders/ {
proxy_set_header X-Served-By $orders_backend; # lets you grep logs per backend
proxy_pass http://$orders_backend;
}
# Everything else is still the monolith, exactly as before.
location / {
proxy_pass http://legacy_erp;
}
}
The split_clients directive, from the nginx documentation, creates a variable for A/B testing: it hashes the string you give it and maps the hash range onto the percentages you list, with a star as the catch-all. Because the hash input here is the client address plus a label, a given client keeps hitting the same backend, which avoids a user seeing different behaviour on each refresh. If you need stickiness per logged-in user rather than per IP, hash a cookie or header instead.
Log the backend name on every request (the X-Served-By header above, or a log_format field). During a ramp you will want to compare error rates for the old and new path side by side, and without a per-backend field that comparison is guesswork.
Worked example: suppose the orders path receives 20,000 requests a day. At 10 percent the new service sees about 2,000 of them, enough to surface a bad edge case within a day while 18,000 requests are still served by the monolith. Ramping to 25 percent gives 5,000, then 50 percent gives 10,000, then 100 percent. The numbers are arithmetic on an assumed volume, not a benchmark, but they show why a percentage ramp finds problems with a bounded blast radius.
Which slice should you extract first?
Pick the first slice for learning, not for glory. You want something whose boundary is already visible, whose data is mostly its own, and whose failure is survivable. Score candidates against these criteria before you touch any code.
Candidate slice
Good first choice when
Warning sign
Read-only reporting or search
It only reads data and a stale result is acceptable for a few seconds
It joins across ten tables owned by different modules
Notifications, email, receipt printing
It is triggered by events and has few inbound dependencies
It reads half the business tables to build its message
Pricing or tax calculation
It is mostly pure logic with a clear input and output
Rules are scattered across stored procedures and UI code
Core ledger or inventory posting
Almost never first; extract it last, if at all
Every other module writes to it inside one transaction
Reporting and notifications teach you the routing, deployment and observability mechanics cheaply. Save the transactional core for when the team has already rehearsed the whole loop several times.
Who owns the data after you extract a slice?
Routing is the easy half. Data is the hard half, because a service that still reads and writes the monolith database is a distributed monolith with extra network hops. The goal is that exactly one system owns each piece of data and everyone else gets it through an API or an event. There are three common ways to get there during the transition.
Approach
How it works
Main risk
Shared database (temporary)
New service reads and writes the monolith tables directly
Schema changes now break two systems; coupling is hidden
Dual write
App writes to the old and new store in the same request
No shared transaction, so one write can succeed and the other fail
Outbox or change data capture
Write once, publish the change from the log or an outbox table
Consumers must be idempotent; the stream adds latency
Dual writes look simple and fail quietly. Illustrative arithmetic: if 20,000 writes a day each have a 0.1 percent chance that the second write fails after the first succeeds, that is 20 diverged records every day, and nothing tells you which ones. The fix is to make one write authoritative and derive the other. The transactional outbox does that: the business row and an event row commit in one local transaction, and a relay publishes the event afterwards. The full Postgres version is in the transactional outbox post.
-- Same transaction: the business row and the event row commit together or not at all.
BEGIN;
UPDATE orders SET status = 'PAID' WHERE id = 8841;
INSERT INTO outbox (aggregate_id, event_type, payload)
VALUES (8841, 'OrderPaid', '{"orderId": 8841, "amount": 250000}');
COMMIT;
-- A relay (a poller, or CDC reading the WAL) publishes outbox rows to the new
-- service afterwards. If the relay dies, the rows are still there. A dual write
-- (DB write, then HTTP call) loses the event the moment the process dies between them.
Change data capture reaches the same result by reading the database log instead of an outbox table, and tools such as Debezium do this for Postgres. Either way the new service builds its own copy of the data from events, and once it is correct you point writes at the new service and stop the old table from being written by the monolith.
Do not leave the shared database in place indefinitely. A service that reads monolith tables behind the facade passes every routing test and still blocks every schema change, so schedule the cut-over of data ownership as an explicit step with a date.
Where do branch by abstraction and an anti-corruption layer fit?
A path-based facade only works when the slice can be reached over HTTP. When the code you are replacing is called in-process from many places, use branch by abstraction, described by Martin Fowler: put an interface in front of the old implementation, move callers onto the interface, build the new implementation behind it, then switch with a flag. The Azure Architecture Center's anti-corruption layer pattern solves the neighbouring problem. It is a translation layer between two subsystems with different models, so the new service's vocabulary does not leak into the legacy system or the reverse. In the sketch below, the adapter is both: the interface carries the branch and the adapter translates.
// Branch by abstraction: callers depend on the interface, never on a backend.
interface OrderPricing {
quote(cartId: string): Promise<Quote>;
}
class LegacyPricing implements OrderPricing {
async quote(cartId: string) {
return legacyPriceCart(cartId); // in-process call into the monolith
}
}
// The anti-corruption layer lives inside this adapter, so the new service's
// vocabulary never leaks into the old code or the other way round.
class PricingServiceClient implements OrderPricing {
async quote(cartId: string) {
const res = await fetch("http://pricing:4002/quotes", {
method: "POST",
body: JSON.stringify({ cartId }),
});
const dto = await res.json();
return {
total: dto.grandTotalMinor / 100, // service speaks minor units
status: dto.state === "FINAL" ? "C" : "P", // monolith still expects "C" / "P"
};
}
}
// The flag decides at runtime, so rolling back is a config flip, not a revert.
export const pricing: OrderPricing = flags.get("pricing.useService")
? new PricingServiceClient()
: new LegacyPricing();
Keep the translation in one adapter. If the monolith expects a single-letter status and the new service returns a word, convert it in that one file, and the day the monolith is retired you delete the adapter and the oddity with it.
How do you roll back, and when do you stop?
Rollback must be a configuration change, not a redeploy. With the facade above, setting the split to 0 percent and reloading nginx sends all traffic back to the monolith in seconds, as long as the monolith still works against the data. That last condition is the catch: once the new service owns writes, a rollback needs the data to flow back, so keep the relay running in both directions until the slice is proven.
Ramp by percentage and hold each step long enough to see a normal business cycle, such as a full trading day for a POS.
Define the rollback trigger before the ramp: an error rate, a latency ceiling or a reconciliation mismatch count.
Delete the old code path only after the new one has carried 100 percent of traffic through at least one full period-end close.
Knowing when to stop matters as much. The goal is not a monolith with zero lines left, it is a system where the remaining core is stable, changes rarely and costs little to run. If the leftover monolith meets that bar, leaving it is a legitimate end state, and a modular monolith is often the better home for what remains.
When should you not migrate at all?
The pattern is a tool for a specific pain, and plenty of teams adopt it without that pain. Treat the following as a checklist, and if two or more apply, fix the monolith first.
The team is small and the real pain is slow releases or flaky tests, which a modular monolith and a better pipeline fix more cheaply.
There is no clear module boundary in the code, so you would be guessing where to cut.
You cannot observe the current system: no per-request logs, metrics or traces, so you could not tell a regression from normal behaviour.
The driver is fashion or a hiring pitch rather than a measured problem such as independent scaling, release isolation or a separate team owning the domain.
Microservices add network failure modes, distributed data and operational load. On a single VPS with one Postgres instance, those costs are real and the benefits often are not.
What is the step-by-step checklist?
Run this loop once per slice. Each pass should be small enough to finish in a few weeks and reversible at every step.
Put the facade in front of the monolith and route 100 percent of traffic through it unchanged. Confirm nothing regresses.
Add per-backend logging and a dashboard before you extract anything.
Choose one slice using the table above and define its API and its data ownership.
Introduce an abstraction or adapter, and build the new service behind it with an anti-corruption layer.
Sync data with an outbox or change data capture, then reconcile old and new until they match.
Ramp the split in steps, watching the rollback trigger, until the slice carries 100 percent.
Move data ownership, delete the old code path, and decide whether another slice is worth it.
The rule to carry: never migrate by replacing, migrate by routing. A facade makes every step reversible, a one-owner-per-table rule keeps the data honest, and a clear stop condition keeps the project from becoming a second rewrite. Extract the slice that teaches you the most, prove it, and ask again whether the next one is worth its cost.