Topic 13.3
Capacity Planning with Real Numbers
In one line
Size a cluster from throughput (MB/s in), replication factor, consumer fan-out, retention and compression: they give you disk, network and broker counts. Add headroom for failures (a lost broker's load moves to others), recovery traffic and growth, then validate with a load test.
Think of it like this
Planning a highway. You count cars per hour at peak, lanes needed, how many lanes must stay open during roadworks, and how much parking the exits need. Guessing leads either to gridlock or to empty concrete.
Key ideas
- 01
Inputs: events/sec, average size, compression ratio, replication factor, number of consumer groups (fan-out), retention, peak-to-average ratio, growth rate.
- 02
Ingest MB/s (compressed) × RF = replication write load; × retention seconds = disk. Outbound = ingest × (RF − 1) for replication + ingest × consumer groups for fetches.
- 03
Per-broker limits come from measured disk throughput, network (for example a 10 Gbps NIC ≈ 1.1 GB/s, keep under ~60–70%), CPU (TLS, compression) and partitions per broker.
- 04
Failure headroom: with N brokers, losing one moves its leadership to N−1; recovery (re-replicating a replacement broker) adds traffic. Plan to run at ≤ 60–70% of limits at peak.
- 05
Validate with
kafka-producer-perf-test/consumer-perf-testagainst a staging cluster sized like production, and monitor for the first months.
Code & diagrams
Inputs: 100,000 events/s x 2 KB = 200 MB/s raw ingest; RF = 3; retention 7 days; 3 consumer groups
Without compression
daily raw: 200 MB/s x 86,400 s = 17.28 TB/day
7-day raw: 17.28 x 7 = 121 TB
replicated storage: 121 x 3 = 363 TB (+ ~20% free-space headroom -> ~435 TB)
replication traffic: 200 x (3 - 1) = 400 MB/s between brokers
consumer egress: 200 x 3 groups = 600 MB/s
total broker egress: ~1,000 MB/s; ingress: 200 + 400 = 600 MB/s
With 3:1 compression (zstd on JSON)
ingest on wire: ~67 MB/s
replicated storage: ~121 TB (+20% -> ~145 TB)
replication traffic: ~133 MB/s; consumer egress ~200 MB/s
Broker count (compressed case), assume per broker: 10 Gbps NIC used to 60% ~ 700 MB/s, 12 TB usable disk
by disk: 145 / 12 ~ 13 brokers
by network: (67 + 133 + 200) / 700 per broker ~ 1 broker -> disk dominates
+ 1 broker failure headroom -> ~15 brokers, or use tiered storage to shrink local disk needsInterview problem
The problem
Capacity calculation
100,000 events/sec, average 2 KB, retention 7 days, replication factor 3. Calculate raw incoming data per second, daily data, seven-day data, replicated storage, network implications and required broker capacity. Then add 3:1 compression and recalculate.
Explain it without notes
Which dimension usually drives Kafka broker count, and how does tiered storage change it?
Practice
Redo the calculation for your own system's traffic.
Trade-offs
- ↔
More headroom costs money but prevents incidents during failures and spikes.
Done when you can
I can calculate storage, network and broker count for a Kafka workload with and without compression.