Command Palette

Search for a command to run...

Hectal
Phase 4Intermediate5 of 18 in Apache Kafka

Replication & Durability

Leaders, followers and the ISR, high watermarks, leader election and unclean elections, the replication factor + acks + min.insync.replicas combination, and surviving an availability-zone failure.

Kafka's durability comes from replicating each partition across brokers and only exposing records once the in-sync replicas have them. Understanding the ISR precisely is the difference between "we use acks=all" and actually knowing what survives a failure.

This phase walks through a broker failure step by step, then designs configurations that survive a broker, then a whole availability zone.

0/4 · 0%
4 topics ~31 min 6 code blocks & diagrams
Start with the first topic
1
4.1

Replication, the ISR, and the High Watermark

Each partition has a leader and followers that fetch from it. Followers caught up within replica.lag.time.max.ms form the in-sync replica set (ISR). The high watermark is the highest offset replicated to all ISR members; consumers only see records below it, so they never read data that could disappear in a failover.

7 min 2 code practice

2
4.2

Leader Election and Unclean Elections

When a partition leader fails, the controller elects a new leader from the ISR, so no committed record is lost. If no ISR member is alive, Kafka waits (unavailable) unless unclean leader election is enabled, which elects an out-of-sync replica and loses data to restore availability.

9 min 1 diagram practice

3
4.3

Durability as a Combination: RF, acks, min.insync, and Friends

No single setting makes Kafka durable. Replication factor sets how many copies exist, acks=all makes producers wait for the ISR, min.insync.replicas sets the minimum copies for a write, unclean election decides what happens when all in-sync copies die, and producer idempotence, retries and the outbox stop loss and duplicates at the edges.

7 min 1 code practice

4
4.4

Rack Awareness and Surviving an AZ Failure

Setting broker.rack to each broker's availability zone makes Kafka spread each partition's replicas across zones, so one zone outage never takes all copies. With RF=3 across 3 AZs and min.insync=2, the cluster keeps accepting durable writes through a full AZ loss. Follower fetching lets consumers read from their own AZ to cut cross-AZ costs.

8 min 1 diagram 1 code practice