Command Palette

Search for a command to run...

Hectal
on-callShift 2.1 · Capacity & Resilience
SEV2

ShopLite p99 3.8s and rising — traffic at 7× normal, unannounced; autoscaling at max tasks

An influencer posted a link

Primary on-call, 20:40 on a Saturday.

load sheddingprioritising critical pathsautoscaling limitsrate limiting abusive traffic

Briefing

A popular creator posted a ShopLite product link to millions of followers. Traffic jumped from 20 to 140 requests per second in four minutes. The service is at its autoscaling maximum (6 tasks), and every request, including checkout, is slow.

Your shift

Decide before you look. Each choice is locked once made; the next decision unlocks after it. There's no perfect path in real incidents either, only better and worse judgement under uncertainty.

20:40· decision 1 of 3

Everything is slow, and checkout is starting to time out.

  • req/s: 140 (normal 20) · /products 88% · /checkout 7% · /search 5%
  • p99: /products 3.8s · /checkout 4.4s · CPU 96% on all 6 tasks

What's your first lever?

Debrief

Under overload, protect the critical path by shedding optional work (caching, disabling features, rate-limiting abusive clients) before adding capacity, because new capacity is slow to arrive and often shifts the bottleneck to something harder to scale. Then automate what you did by hand.

The postmortem

Blameless postmortem · draft

Summary. Unannounced 7× traffic surge from a social media post saturated application CPU; mitigated with CDN caching, feature switches, and edge rate limiting.

Impact. Checkout p99 above 1 s for 12 minutes; ~3% of checkouts timed out.

Timeline

  • 20:36Traffic surge begins
  • 20:40Latency SLO page
  • 20:44Degradation switches enabled; checkout recovers
  • 20:58Edge rate limit for scrapers

Root cause. Traffic exceeded application capacity at the autoscaling limit; product pages were uncached and shared capacity with checkout.

Contributing factors

  • Product pages not cached at the edge at all.
  • Degradation switches existed but were manual only.

Action items

  • preventCache product pages at CDN with short TTL and stale-while-revalidateshoplite
  • mitigateAutomatic degradation when CPU > 85% for 2 minplatform
  • preventWAF rate-based rules on by defaultsre

The practice behind it

Graceful degradation and load shedding

When demand exceeds capacity, something must give. Graceful degradation decides in advance WHAT gives: optional features first, stale data before no data, bots before humans, browsing before checkout. Load shedding rejects excess work early and cheaply (429/503 with Retry-After) instead of letting everything slow down until everything fails.

Your turn

01

Rank ShopLite's features from 'shed first' to 'protect at all costs'.

Interview questions

01

Your service is overloaded and autoscaling is maxed out. What do you do?

0/4 · 0%