Command Palette

Search for a command to run...

PHASE 14Advanced ~15 min· topic 9 of 11

Topic 14.9

Deployment Strategies: Rolling, Blue-Green, Canary & Rollback

In one line

Most outages start with a change, so how you release matters as much as what you release. Rolling deploys replace instances a few at a time. Blue-green runs the new version beside the old and switches traffic at once (and back if needed). Canary releases send a small share of traffic to the new version and widen it only while metrics stay healthy. Pair every strategy with health checks, automatic rollback, and feature flags for risky behaviour.

0/11 · 0%

Think of it like this

A restaurant changing its biryani recipe. Rolling: one cook at a time switches. Blue-green: a second kitchen cooks the new recipe and the waiters switch kitchens at once. Canary: serve the new recipe to one table in twenty and watch faces before telling everyone.

Words you'll meet

New words in this topic, in plain English. Come back here whenever one feels fuzzy.

Rolling deploy
Replacing instances of a service with the new version in small batches.
Blue-green deploy
Running two full environments and switching traffic from the old (blue) to the new (green).
Canary release
Sending a small share of traffic to the new version first and widening it while it stays healthy.
Rollback
Returning to the previous version after a bad release.
Bake time
How long a release step runs before moving on, to let problems show up.
Progressive delivery
Releasing changes gradually with automated checks, using canaries and feature flags.

Step by step

01Three ways to ship the same change

Tiffin ships a new pricing engine. Each strategy trades cost, speed, and how many users a bad release can hurt.

Three ways to ship the same changediagram
Rendering diagram…

02A canary with automatic analysis

With Argo Rollouts, each step waits, then a metric query decides whether to continue. If the canary's error rate exceeds 1%, the rollout aborts and traffic returns to the stable version.

rollout.yaml (excerpt)whole fileyaml
strategy:
  canary:
    steps:
      - setWeight: 5
      - pause: { duration: 10m }
      - analysis: { templates: [{ templateName: error-rate }] }
      - setWeight: 25
      - pause: { duration: 10m }
      - analysis: { templates: [{ templateName: error-rate }] }
      - setWeight: 50
      - pause: { duration: 15m }
---
# AnalysisTemplate error-rate: fail if 5xx share of canary traffic > 1% (checked every minute, 3 failures abort)
terminal
$ kubectl argo rollouts get rollout pricing -n tiffin | head -8
── expected output ──
Name: pricing
Status: ✖ Degraded
Message: RolloutAborted: metric "error-rate" assessed Failed due to failed (3) > failureLimit (2)
Strategy: Canary
Step: 2/8
SetWeight: 0
ActualWeight: 0
Images: tiffin/pricing:1.41.0 (stable)
Only 5% of traffic saw the bad version, for about 3 minutes.

03Blue-green for a risky infrastructure change

For a runtime upgrade (Java 21 to 25) that changes memory behaviour, the team uses blue-green: green is deployed and load-tested with synthetic traffic, then the load balancer switches. Blue stays running for an hour so switching back takes seconds.

terminal
$ aws elbv2 modify-listener --listener-arn $LISTENER --default-actions Type=forward,TargetGroupArn=$GREEN_TG
aws elbv2 describe-target-health --target-group-arn $GREEN_TG --query 'TargetHealthDescriptions[].TargetHealth.State'
── expected output ──
[
"healthy",
"healthy",
"healthy",
"healthy"
]

Break it on purpose

Errors are the best teachers. Make each change, read the error, guess what went wrong, then reveal the answer.

Break #1

A canary measured on the wrong traffic

The canary gets 5% of requests, but analysis compares the overall error rate (both versions mixed) with a threshold.

terminal
$ # metrics during the rollout
── what you'll see ──
overall 5xx: 0.6% (under the 1% threshold) → promoted to 100%
canary-only 5xx was 12%
# after full rollout: 12% of all orders fail

Myth vs fact

Myth

If it passed staging, it's safe to release to everyone at once.

Fact

Production has real traffic patterns, data, and scale that staging doesn't. Gradual rollout with monitoring catches what staging can't.

Pro corner

Extra depth for experienced readers. New to this? Skip it for now and come back later.

  • ▸

    Track change failure rate and mean time to restore (DORA metrics). Fast, automated rollback usually improves reliability more than slower, more manual release processes.

Remember this

  1. 1

    Rolling: replace instances in batches (Kubernetes default). Cheap and simple, but old and new run together for a while, and rolling back is another rolling deploy.

  2. 2

    Blue-green: deploy the new version (green) fully, test it, switch the load balancer from blue to green, and keep blue ready to switch back instantly. Needs double capacity during the switch, and the database must work with both.

  3. 3

    Canary: route 1-5% of traffic to the new version, compare error rate and latency with the old version, then step up (10%, 25%, 50%, 100%). Problems hit few users and are caught by metrics.

  4. 4

    Automated gates: each step checks health metrics against thresholds and rolls back automatically on breach. Human approval is a backstop, not the main safety.

  5. 5

    Feature flags separate deploying code from releasing behaviour: code ships dark and is turned on gradually, and turned off in seconds without a redeploy.

  6. 6

    Database changes must be backward-compatible (expand/contract, Topic 11.17), or no deployment strategy can roll back safely.

Explain it without notes

01

When is blue-green better than canary?

02

Why can't a bad database migration be rolled back by switching traffic?

Practice

01

Write the promotion and rollback rules for a canary.

02

Plan a rollback drill for your service.

Trade-offs

  • ↔

    Rolling is cheapest but slow to roll back. Blue-green rolls back instantly but needs double capacity. Canary limits the blast radius but needs good per-version metrics and takes longer to reach 100%.

Run it in production

Done when you can

  • I can compare rolling, blue-green, and canary releases.

  • My releases have automated health gates and rollback.

  • I keep schema changes backward-compatible.

Back to phase