Topic 10.3
Broker Resources: JVM, GC, Disks, File Handles, Network
In one line
A broker needs a right-sized heap with a low-pause collector, lots of free RAM for page cache, fast disks (several log dirs or JBOD), high file-descriptor and memory-map limits, and network bandwidth sized for produce, replication and consumer fan-out combined.
Think of it like this
Running a busy warehouse. You need enough staff (CPU threads), a clear loading dock (network), fast forklifts (disks), a large staging area (page cache), and a supervisor who doesn't freeze for minutes during shift changes (GC pauses).
Key ideas
- 01
JVM: heap 4–8 GB is typical; G1GC by default; watch GC pause times (long pauses stall request handling and can drop brokers from the ISR or controller sessions). Kafka 4.0 brokers need Java 17+.
- 02
Threads:
num.network.threads(accept and read requests),num.io.threads(process requests, disk I/O),num.replica.fetchers(follower fetch threads). MetricsNetworkProcessorAvgIdlePercentandRequestHandlerAvgIdlePercentbelow ~30% indicate saturation. - 03
OS limits: open files (each segment's log, index and timeindex, plus sockets) often need
ulimit -n 100000or more;vm.max_map_counthigh enough for memory-mapped indexes; avoid swap (vm.swappiness=1). - 04
Disks: SSDs or fast cloud volumes (throughput-provisioned), XFS filesystem, multiple
log.dirsto spread I/O. Watch disk utilisation and latency; a slow disk on one broker slows every partition it leads and its followers' ISR membership. - 05
Network: per broker, bytes in = produce + replication in; bytes out = replication out + consumer fetches (× number of consumer groups). A topic with 5 consumer groups multiplies outbound traffic by 5.
Code & diagrams
# /etc/security/limits.d/kafka.conf
kafka soft nofile 128000
kafka hard nofile 128000
# sysctl
vm.max_map_count=262144
vm.swappiness=1
# JVM (kafka-server-start.sh reads KAFKA_HEAP_OPTS / KAFKA_JVM_PERFORMANCE_OPTS)
export KAFKA_HEAP_OPTS="-Xms6g -Xmx6g"
export KAFKA_JVM_PERFORMANCE_OPTS="-XX:+UseG1GC -XX:MaxGCPauseMillis=20 -XX:InitiatingHeapOccupancyPercent=35"
# server.properties
num.network.threads=8
num.io.threads=16
num.replica.fetchers=4
log.dirs=/data1/kafka,/data2/kafkaWhen it breaks
Broker hits the open-file limit
What you see
Too many open files errors; the broker can't roll segments or accept connections and may shut down, taking partitions offline.
Fix & prevent
Raise nofile limits well above partitions × files per segment plus connections; monitor open file descriptors.
Explain it without notes
How do consumer groups affect broker network sizing?
Practice
Check your lab broker's JVM heap, GC collector and open file count with jcmd and lsof.
Trade-offs
- ↔
Bigger heaps rarely help Kafka and increase GC pause risk; RAM is better spent on page cache.
Done when you can
I can size and tune broker JVM, threads, OS limits, disks and network.