Topic 14.10
Limiting Blast Radius: Cells, Shuffle Sharding & Regions
In one line
Failures will happen, so design how far each one can spread. Split the system into independent cells, each a complete copy of the stack serving a subset of customers, so one cell's bad deploy or overload affects only its customers. Shuffle sharding gives each customer a random small set of servers, so one bad customer can't take down everyone. Deploy region by region and cell by cell, and avoid global dependencies that every cell shares.
Think of it like this
A large ship divided into watertight compartments. A hole floods one compartment, not the whole ship. Cells are compartments for software.
Words you'll meet
New words in this topic, in plain English. Come back here whenever one feels fuzzy.
- Blast radius
- The scope of impact of a single failure: how many users, tenants, or features are affected.
- Cell
- An independent, complete copy of a service stack serving a subset of customers.
- Cell router
- The thin layer that sends each request to the customer's cell.
- Shuffle sharding
- Assigning each customer a random small subset of resources to make overlap between customers rare.
- Static stability
- A system that keeps working with its last known state when a dependency (like a control plane) fails.
Step by step
01From one big stack to cells
Tiffin's B2B lunch platform serves 2,000 corporate customers from one stack. A bad query from one large customer once took the shared database down for everybody. They split into 8 cells, each with its own database, cache, and app servers, and put the largest customers in their own cells.
02Shuffle sharding the API workers
Inside a cell, 8 worker pools serve requests. Each tenant is hashed to 2 of the 8. With 28 possible pairs, a tenant whose requests crash workers takes down its own 2 pools, and only tenants sharing exactly that pair lose everything. Everyone else still has at least one healthy pool.
// Deterministically pick k of n pools for a tenant.
List<Integer> poolsFor(String tenant, int n, int k) {
var rnd = new java.util.Random(murmur3(tenant)); // stable seed per tenant
var all = new ArrayList<Integer>();
for (int i = 0; i < n; i++) all.add(i);
java.util.Collections.shuffle(all, rnd);
return all.subList(0, k); // e.g. [5, 2]
}
// 8 pools, 2 per tenant: 28 combinations, so a given pair is shared by about 1/28 of tenants03Releasing cell by cell
Deploys go to a 'canary cell' with internal and friendly customers first, bake for an hour, then two cells, then the rest. A bad release has never affected more than one cell since the change.
Break it on purpose
Errors are the best teachers. Make each change, read the error, guess what went wrong, then reveal the answer.
Break #1
Cells that share a hidden dependency
The 8 cells are independent, except they all read feature flags from one global flag service at startup and on every request.
Myth vs fact
Myth
More redundancy always means a smaller blast radius.
Fact
Redundant copies of the same shared system can fail together (a bad config or deploy hits all of them). Isolation means independent copies, changed one at a time.
Pro corner
Extra depth for experienced readers. New to this? Skip it for now and come back later.
- ▸
Keep the cell router extremely simple (a cached mapping and forwarding), so it rarely changes and rarely fails. Moving a tenant between cells is then a data migration plus a mapping change, which is also how you rebalance load.
Remember this
- 1
Blast radius = how many users or features one failure affects. A single shared database, cache, or deploy gives a blast radius of 'everyone'.
- 2
Cells: complete, independent stacks (app, database, cache, queue), each serving a fixed set of tenants or users. A thin routing layer maps each customer to their cell. A cell failure hits only its share (for example 1/10 of customers).
- 3
Shuffle sharding: give each customer a random combination of a few servers out of many (for example 2 of 8). Two customers rarely share the exact same combination, so a 'poison' customer that crashes its servers leaves almost every other customer with at least one healthy server.
- 4
Deploy by blast radius: release to one cell (or one availability zone, one region) first, bake, then the next. A bad release stops at the first cell.
- 5
Avoid global single points: shared control planes, DNS, auth, configuration, and the routing layer itself must be simple, redundant, and changed carefully, since they touch every cell.
- 6
Regions: multi-region active-active gives the largest isolation (and lowest latency for users), at the cost of data replication and conflict handling. Many systems keep data per region (users homed in one region).
Explain it without notes
Why does shuffle sharding work better than plain sharding against a poison tenant?
What's the cost of a cell architecture?
Practice
Compute the blast radius for 10 cells, and for shuffle sharding with 8 pools and 2 per tenant.
Identify the global dependencies in a system you know.
Trade-offs
- ↔
Cells and shuffle sharding cap the damage from failures and bad tenants, but add cost and operational complexity. For small systems, availability zones and careful deploys may give enough isolation.
Run it in production
You've designed it. Now build, operate, and break the same idea hands-on in the DevOps courses:
Done when you can
I can explain blast radius and cells.
I can apply shuffle sharding.
I deploy cell by cell and make shared dependencies statically stable.