Command Palette

Search for a command to run...

Hectal
PHASE 10Advanced ~7 min· topic 3 of 6

Topic 10.3

Broker Resources: JVM, GC, Disks, File Handles, Network

In one line

A broker needs a right-sized heap with a low-pause collector, lots of free RAM for page cache, fast disks (several log dirs or JBOD), high file-descriptor and memory-map limits, and network bandwidth sized for produce, replication and consumer fan-out combined.

0/6 · 0%

Think of it like this

Running a busy warehouse. You need enough staff (CPU threads), a clear loading dock (network), fast forklifts (disks), a large staging area (page cache), and a supervisor who doesn't freeze for minutes during shift changes (GC pauses).

Key ideas

  1. 01

    JVM: heap 4–8 GB is typical; G1GC by default; watch GC pause times (long pauses stall request handling and can drop brokers from the ISR or controller sessions). Kafka 4.0 brokers need Java 17+.

  2. 02

    Threads: num.network.threads (accept and read requests), num.io.threads (process requests, disk I/O), num.replica.fetchers (follower fetch threads). Metrics NetworkProcessorAvgIdlePercent and RequestHandlerAvgIdlePercent below ~30% indicate saturation.

  3. 03

    OS limits: open files (each segment's log, index and timeindex, plus sockets) often need ulimit -n 100000 or more; vm.max_map_count high enough for memory-mapped indexes; avoid swap (vm.swappiness=1).

  4. 04

    Disks: SSDs or fast cloud volumes (throughput-provisioned), XFS filesystem, multiple log.dirs to spread I/O. Watch disk utilisation and latency; a slow disk on one broker slows every partition it leads and its followers' ISR membership.

  5. 05

    Network: per broker, bytes in = produce + replication in; bytes out = replication out + consumer fetches (× number of consumer groups). A topic with 5 consumer groups multiplies outbound traffic by 5.

Code & diagrams

broker-tuning.shbash
# /etc/security/limits.d/kafka.conf
kafka  soft  nofile  128000
kafka  hard  nofile  128000
# sysctl
vm.max_map_count=262144
vm.swappiness=1
# JVM (kafka-server-start.sh reads KAFKA_HEAP_OPTS / KAFKA_JVM_PERFORMANCE_OPTS)
export KAFKA_HEAP_OPTS="-Xms6g -Xmx6g"
export KAFKA_JVM_PERFORMANCE_OPTS="-XX:+UseG1GC -XX:MaxGCPauseMillis=20 -XX:InitiatingHeapOccupancyPercent=35"
# server.properties
num.network.threads=8
num.io.threads=16
num.replica.fetchers=4
log.dirs=/data1/kafka,/data2/kafka

When it breaks

Broker hits the open-file limit

What you see

Too many open files errors; the broker can't roll segments or accept connections and may shut down, taking partitions offline.

Fix & prevent

Raise nofile limits well above partitions × files per segment plus connections; monitor open file descriptors.

Explain it without notes

01

How do consumer groups affect broker network sizing?

Practice

01

Check your lab broker's JVM heap, GC collector and open file count with jcmd and lsof.

Trade-offs

  • ↔

    Bigger heaps rarely help Kafka and increase GC pause risk; RAM is better spent on page cache.

Done when you can

  • I can size and tune broker JVM, threads, OS limits, disks and network.