Command Palette

Search for a command to run...

PHASE 11Intermediate ~15 min· topic 16 of 16

Topic 11.17

Zero-Downtime Migrations: Expand/Contract & the Strangler Fig

In one line

Changing a live system without an outage needs small, backward-compatible steps. For databases, use expand/contract: add the new column or table, write to both, backfill, switch reads, then remove the old one, with each step deployable and reversible on its own. For replacing a whole system, use the strangler fig: put a routing layer in front, move one route or feature at a time to the new service, and retire the old system when nothing is left.

0/16 · 0%

Think of it like this

Renovating a busy railway station. You don't close it for a year. You build the new platform beside the old one, run trains on both for a while, move passengers over line by line, and only then demolish the old platform.

Words you'll meet

New words in this topic, in plain English. Come back here whenever one feels fuzzy.

Expand/contract
Changing a schema by first adding the new structure alongside the old, then removing the old one after code has moved.
Backfill
Copying or computing data for existing rows into a new column or table.
Dual write
Writing the same data to the old and new places during a migration.
Strangler fig
Replacing a legacy system gradually by routing features one by one to a new system.
Shadow traffic
Sending a copy of real requests to the new system to compare results without affecting users.
Online DDL
Schema changes that don't block reads and writes for long, such as CREATE INDEX CONCURRENTLY.

Step by step

01Renaming a column safely

Tiffin wants to replace orders.address (free text) with delivery_address_id pointing to a proper addresses table. Old app instances still read address during a deploy, so the change happens in stages over several releases.

terminal
$ psql -c "ALTER TABLE orders ADD COLUMN delivery_address_id BIGINT REFERENCES addresses(id)"
psql -c "CREATE INDEX CONCURRENTLY orders_delivery_address_idx ON orders (delivery_address_id)"
── expected output ──
ALTER TABLE
CREATE INDEX
Adding a nullable column is instant. CONCURRENTLY builds the index without blocking writes.
Renaming a column safelydiagram
Rendering diagram…

02Backfilling in batches

A job converts old addresses 5,000 rows at a time, pausing between batches and watching replica lag, so production traffic isn't affected.

backfill.sql (run in a loop by a job)whole filesql
WITH batch AS (
  SELECT id, address FROM orders
  WHERE delivery_address_id IS NULL
  ORDER BY id
  LIMIT 5000
  FOR UPDATE SKIP LOCKED
)
UPDATE orders o
SET delivery_address_id = upsert_address(b.address)   -- helper returns an address id
FROM batch b
WHERE o.id = b.id;
-- repeat until 0 rows updated; sleep 200 ms between batches; stop if replica lag > 5 s

03Strangling the old dispatch system

Tiffin's old dispatch is part of a big monolith. A new dispatch service is built, and the gateway routes one city at a time to it: Guwahati first (smaller), comparing assignments with the monolith's in shadow mode for a week, then Pune, then Mumbai. When the last city moves, the monolith's dispatch code is deleted.

Strangling the old dispatch systemdiagram
Rendering diagram…

Break it on purpose

Errors are the best teachers. Make each change, read the error, guess what went wrong, then reveal the answer.

Break #1

Dropping a column in the same deploy

A release removes orders.address from the code and drops the column in a migration that runs at deploy start.

terminal
$ # during the rolling deploy
── what you'll see ──
old pods (still running): ERROR: column "address" does not exist
5xx on order creation: 48% for 6 minutes

Myth vs fact

Myth

Big-bang rewrites are faster than gradual migration.

Fact

They usually take longer, freeze features for months, and fail at cut-over. Gradual migration delivers value early and keeps rollback possible at every step.

Pro corner

Extra depth for experienced readers. New to this? Skip it for now and come back later.

  • ▸

    Use lock_timeout (for example SET lock_timeout = '3s') before schema changes, so a migration waiting for a lock fails quickly instead of queuing behind a long query and blocking every other query behind itself.

Remember this

  1. 1

    During a rolling deploy, old and new code run at the same time against the same database, so every schema change must work with both versions.

  2. 2

    Expand/contract for a column rename: (1) add the new column, (2) deploy code writing both, (3) backfill old rows in batches, (4) switch reads to the new column, (5) stop writing the old one, (6) drop it. Never rename or drop in one step.

  3. 3

    Backfills run in small batches with pauses, so they don't lock tables or flood replicas. Long-running locks (adding a column with a default on old databases, building indexes) need online methods (CREATE INDEX CONCURRENTLY).

  4. 4

    Strangler fig: route traffic through a facade (gateway or proxy). Move one capability at a time to the new service, starting with something low-risk, while the rest still goes to the old system. Compare results (shadow traffic) before switching.

  5. 5

    Data during the transition: either the old system stays the source of truth and the new one gets a copy (CDC), or you dual-write with reconciliation. Pick one owner per piece of data at each stage.

  6. 6

    Every step needs a rollback plan, and the migration isn't done until the old path is deleted.

Explain it without notes

01

Why can't a migration both add a new column and remove the old one in one release?

02

How do you know when a strangler-fig migration is finished?

Practice

01

Plan an expand/contract migration to split users.name into first and last name.

02

Plan a strangler-fig migration for one feature of a monolith you know.

Trade-offs

  • ↔

    Gradual migrations take more releases and temporary complexity (dual writes, routing rules) but avoid outages and keep rollback possible. Big-bang changes look simpler but concentrate all the risk at one moment.

Run it in production

Done when you can

  • I split schema changes into expand and contract steps.

  • I backfill in batches with online index builds.

  • I can plan a strangler-fig migration with rollback at each step.

Back to phase