Topic 11.17
Zero-Downtime Migrations: Expand/Contract & the Strangler Fig
In one line
Changing a live system without an outage needs small, backward-compatible steps. For databases, use expand/contract: add the new column or table, write to both, backfill, switch reads, then remove the old one, with each step deployable and reversible on its own. For replacing a whole system, use the strangler fig: put a routing layer in front, move one route or feature at a time to the new service, and retire the old system when nothing is left.
Think of it like this
Renovating a busy railway station. You don't close it for a year. You build the new platform beside the old one, run trains on both for a while, move passengers over line by line, and only then demolish the old platform.
Words you'll meet
New words in this topic, in plain English. Come back here whenever one feels fuzzy.
- Expand/contract
- Changing a schema by first adding the new structure alongside the old, then removing the old one after code has moved.
- Backfill
- Copying or computing data for existing rows into a new column or table.
- Dual write
- Writing the same data to the old and new places during a migration.
- Strangler fig
- Replacing a legacy system gradually by routing features one by one to a new system.
- Shadow traffic
- Sending a copy of real requests to the new system to compare results without affecting users.
- Online DDL
- Schema changes that don't block reads and writes for long, such as
CREATE INDEX CONCURRENTLY.
Step by step
01Renaming a column safely
Tiffin wants to replace orders.address (free text) with delivery_address_id pointing to a proper addresses table. Old app instances still read address during a deploy, so the change happens in stages over several releases.
02Backfilling in batches
A job converts old addresses 5,000 rows at a time, pausing between batches and watching replica lag, so production traffic isn't affected.
WITH batch AS (
SELECT id, address FROM orders
WHERE delivery_address_id IS NULL
ORDER BY id
LIMIT 5000
FOR UPDATE SKIP LOCKED
)
UPDATE orders o
SET delivery_address_id = upsert_address(b.address) -- helper returns an address id
FROM batch b
WHERE o.id = b.id;
-- repeat until 0 rows updated; sleep 200 ms between batches; stop if replica lag > 5 s03Strangling the old dispatch system
Tiffin's old dispatch is part of a big monolith. A new dispatch service is built, and the gateway routes one city at a time to it: Guwahati first (smaller), comparing assignments with the monolith's in shadow mode for a week, then Pune, then Mumbai. When the last city moves, the monolith's dispatch code is deleted.
Break it on purpose
Errors are the best teachers. Make each change, read the error, guess what went wrong, then reveal the answer.
Break #1
Dropping a column in the same deploy
A release removes orders.address from the code and drops the column in a migration that runs at deploy start.
Myth vs fact
Myth
Big-bang rewrites are faster than gradual migration.
Fact
They usually take longer, freeze features for months, and fail at cut-over. Gradual migration delivers value early and keeps rollback possible at every step.
Pro corner
Extra depth for experienced readers. New to this? Skip it for now and come back later.
- ▸
Use
lock_timeout(for exampleSET lock_timeout = '3s') before schema changes, so a migration waiting for a lock fails quickly instead of queuing behind a long query and blocking every other query behind itself.
Remember this
- 1
During a rolling deploy, old and new code run at the same time against the same database, so every schema change must work with both versions.
- 2
Expand/contract for a column rename: (1) add the new column, (2) deploy code writing both, (3) backfill old rows in batches, (4) switch reads to the new column, (5) stop writing the old one, (6) drop it. Never rename or drop in one step.
- 3
Backfills run in small batches with pauses, so they don't lock tables or flood replicas. Long-running locks (adding a column with a default on old databases, building indexes) need online methods (
CREATE INDEX CONCURRENTLY). - 4
Strangler fig: route traffic through a facade (gateway or proxy). Move one capability at a time to the new service, starting with something low-risk, while the rest still goes to the old system. Compare results (shadow traffic) before switching.
- 5
Data during the transition: either the old system stays the source of truth and the new one gets a copy (CDC), or you dual-write with reconciliation. Pick one owner per piece of data at each stage.
- 6
Every step needs a rollback plan, and the migration isn't done until the old path is deleted.
Explain it without notes
Why can't a migration both add a new column and remove the old one in one release?
How do you know when a strangler-fig migration is finished?
Practice
Plan an expand/contract migration to split users.name into first and last name.
Plan a strangler-fig migration for one feature of a monolith you know.
Trade-offs
- ↔
Gradual migrations take more releases and temporary complexity (dual writes, routing rules) but avoid outages and keep rollback possible. Big-bang changes look simpler but concentrate all the risk at one moment.
Run it in production
You've designed it. Now build, operate, and break the same idea hands-on in the DevOps courses:
Done when you can
I split schema changes into expand and contract steps.
I backfill in batches with online index builds.
I can plan a strangler-fig migration with rollback at each step.