Command Palette

Search for a command to run...

Hectal
PHASE 14Advanced ~9 min· topic 5 of 5

Topic 14.5

Failure Injection: Breaking Kafka on Purpose

In one line

Rehearse fifteen failures (producer network loss, broker crash, leader crash, lagging follower, consumer crash before commit, slow processing, rebalance storms, hot partition, full disk, network partition, incompatible schema, poison message, slow database, region failure, massive backlog) and answer ten questions for each.

0/5 · 0%

Think of it like this

Flight simulators train pilots on engine failures before they ever meet one in the air. Kafka game days do the same for your team and your code.

Key ideas

  1. 01

    The ten questions: what breaks; what continues; what can be lost; what can be duplicated; what happens to ordering; what happens to lag; how recovery works; which configuration matters; which monitoring catches it; how to prevent recurrence.

  2. 02

    Tools: docker kill/pause, tc netem for latency and loss, iptables or Kubernetes NetworkPolicies for partitions, fallocate to fill disks, Toxiproxy between clients and brokers, Chaos Mesh or AWS FIS in real environments.

  3. 03

    Run in staging first with a hypothesis and a stop condition; automate the scenarios you care about most and repeat them after major changes.

Code & diagrams

scenarios.txttext
Scenario                          Lost?                         Duplicated?                   Key config / defence
A Producer loses network          no (if app retries/outbox)    no (idempotence)              delivery.timeout.ms, outbox
B Broker crashes                  no (acks=all, min.isr=2)      no (idempotent retries)       RF=3, min.insync=2
C Leader crashes                  unacked only                  no                            clean election only
D Follower falls behind           no                            no                            replica.lag.time.max.ms, URP alert
E Consumer dies before commit     no                            yes, reprocessed              idempotent consumer / inbox
F Processing > poll interval      no                            yes, reprocessed              max.poll.interval.ms, pause/resume
G Constant rebalances             no                            yes                           cooperative/KIP-848, static membership
H Hot partition                   no                            no                            key design, bucketing
I Disk full                       writes rejected on that dir   no                            disk alerts, retention, tiered storage
J Network partition               no (min.isr blocks writes)    possible on retries           rack awareness, min.insync
K Incompatible schema deploy      no, but consumers fail        no                            registry compatibility in CI
L Poison message                  no                            no                            DLT, error-handling deserializer
M Slow downstream DB              no                            maybe (timeouts)              backpressure, pause, batching
N Region failure                  un-mirrored tail (RPO)        yes after failover            MM2, offset sync, idempotency
O Massive backlog                 if lag > retention            no                            time-lag alerts, scaling, retention
chaos.shbash
# B/C: kill the leader of partition 0 while producing
LEADER=$(kafka-topics.sh --bootstrap-server $B --describe --topic payments | awk '/Partition: 0/ {print $6}')
docker kill kafka$LEADER

# J: partition broker 3 from the others (lab)
docker network disconnect kafka-lab_default kafka3

# I: fill the disk of broker 2
docker exec kafka2 fallocate -l 18G /var/lib/kafka/data/fill.bin

# E: consumer dies after processing, before commit (kill -9 mid-batch)
pkill -9 -f payment-consumer

Interview problem

The problem

Game day for the order platform

Plan a game day for an order platform (outbox → Kafka → payment, inventory, notification consumers). Choose five scenarios, state hypotheses, and define what you'll measure.

Explain it without notes

01

List the ten questions to answer for each Kafka failure scenario.

Practice

01

Run scenarios B, E and L in your lab and fill in the ten answers for each.

Trade-offs

  • ↔

    Game days take effort but turn unknown failure behaviour into documented, tested expectations.

Done when you can

  • I can inject and analyse all fifteen Kafka failure scenarios.