Topic 14.5
Failure Injection: Breaking Kafka on Purpose
In one line
Rehearse fifteen failures (producer network loss, broker crash, leader crash, lagging follower, consumer crash before commit, slow processing, rebalance storms, hot partition, full disk, network partition, incompatible schema, poison message, slow database, region failure, massive backlog) and answer ten questions for each.
Think of it like this
Flight simulators train pilots on engine failures before they ever meet one in the air. Kafka game days do the same for your team and your code.
Key ideas
- 01
The ten questions: what breaks; what continues; what can be lost; what can be duplicated; what happens to ordering; what happens to lag; how recovery works; which configuration matters; which monitoring catches it; how to prevent recurrence.
- 02
Tools:
docker kill/pause,tc netemfor latency and loss, iptables or Kubernetes NetworkPolicies for partitions,fallocateto fill disks, Toxiproxy between clients and brokers, Chaos Mesh or AWS FIS in real environments. - 03
Run in staging first with a hypothesis and a stop condition; automate the scenarios you care about most and repeat them after major changes.
Code & diagrams
Scenario Lost? Duplicated? Key config / defence
A Producer loses network no (if app retries/outbox) no (idempotence) delivery.timeout.ms, outbox
B Broker crashes no (acks=all, min.isr=2) no (idempotent retries) RF=3, min.insync=2
C Leader crashes unacked only no clean election only
D Follower falls behind no no replica.lag.time.max.ms, URP alert
E Consumer dies before commit no yes, reprocessed idempotent consumer / inbox
F Processing > poll interval no yes, reprocessed max.poll.interval.ms, pause/resume
G Constant rebalances no yes cooperative/KIP-848, static membership
H Hot partition no no key design, bucketing
I Disk full writes rejected on that dir no disk alerts, retention, tiered storage
J Network partition no (min.isr blocks writes) possible on retries rack awareness, min.insync
K Incompatible schema deploy no, but consumers fail no registry compatibility in CI
L Poison message no no DLT, error-handling deserializer
M Slow downstream DB no maybe (timeouts) backpressure, pause, batching
N Region failure un-mirrored tail (RPO) yes after failover MM2, offset sync, idempotency
O Massive backlog if lag > retention no time-lag alerts, scaling, retention# B/C: kill the leader of partition 0 while producing
LEADER=$(kafka-topics.sh --bootstrap-server $B --describe --topic payments | awk '/Partition: 0/ {print $6}')
docker kill kafka$LEADER
# J: partition broker 3 from the others (lab)
docker network disconnect kafka-lab_default kafka3
# I: fill the disk of broker 2
docker exec kafka2 fallocate -l 18G /var/lib/kafka/data/fill.bin
# E: consumer dies after processing, before commit (kill -9 mid-batch)
pkill -9 -f payment-consumerInterview problem
The problem
Game day for the order platform
Plan a game day for an order platform (outbox → Kafka → payment, inventory, notification consumers). Choose five scenarios, state hypotheses, and define what you'll measure.
Explain it without notes
List the ten questions to answer for each Kafka failure scenario.
Practice
Run scenarios B, E and L in your lab and fill in the ten answers for each.
Trade-offs
- ↔
Game days take effort but turn unknown failure behaviour into documented, tested expectations.
Done when you can
I can inject and analyse all fifteen Kafka failure scenarios.