Command Palette

Search for a command to run...

Hectal
PHASE 2Intermediate ~7 min· topic 2 of 4

Topic 2.2

Batching, linger.ms, and Compression

In one line

Producers send records in per-partition batches; batch.size caps a batch's bytes and linger.ms waits a little for it to fill. Bigger batches compress better and cut requests, raising throughput at a small latency cost. Compression (lz4, zstd, snappy, gzip) trades CPU for network and disk.

0/4 · 0%

Think of it like this

A delivery van. Leaving the moment one parcel is loaded is fast for that parcel but wasteful; waiting a few minutes to fill the van moves far more parcels per trip.

Key ideas

  1. 01

    batch.size (default 16 KB) is the maximum bytes per partition batch; linger.ms is how long to wait for more records. Kafka 4.0 raised the default linger.ms from 0 to 5 ms because the tiny delay improves batching for most workloads. High-throughput producers often use 64–256 KB batches and 5–20 ms linger.

  2. 02

    Batches are compressed as a unit by the producer, stored compressed on the broker (if the topic's compression.type is producer or matches), replicated compressed, and decompressed by consumers. So compression saves network, disk and replication bandwidth.

  3. 03

    Codecs: lz4 and snappy are fast with moderate ratios; zstd gives the best ratio at reasonable CPU (a common default today); gzip compresses well but costs more CPU. JSON typically compresses 3–10×.

  4. 04

    Bigger batches compress better (more repetition), so batching and compression reinforce each other. Keyed traffic spread over many partitions forms smaller batches per partition, one reason not to over-partition.

  5. 05

    Watch producer metrics: batch-size-avg, records-per-request-avg, compression-rate-avg, record-queue-time-avg (time waiting in the accumulator) and request-latency-avg.

Code & diagrams

throughput-producer.propertiesproperties
acks=all
enable.idempotence=true
compression.type=zstd
batch.size=131072          # 128 KB per partition batch
linger.ms=10               # wait up to 10 ms to fill batches
buffer.memory=67108864     # 64 MB accumulator
max.in.flight.requests.per.connection=5
compare.shbash
for c in none lz4 zstd gzip; do
  kafka-producer-perf-test.sh --topic perf --num-records 1000000 --record-size 1024 --throughput -1 \
    --payload-file sample-json-events.txt \
    --producer-props bootstrap.servers=$B compression.type=$c linger.ms=10 batch.size=131072 | tail -1
done
# none: 48.2 MB/s sent, network out 48.2 MB/s
# lz4:  95.7 MB/s sent, network out ~22 MB/s
# zstd: 88.1 MB/s sent, network out ~12 MB/s
# gzip: 41.5 MB/s sent, network out ~11 MB/s   (illustrative; measure your data)

When it breaks

linger.ms=0 with many partitions and small records

What you see

Thousands of tiny requests per second; brokers spend CPU on request overhead; compression ratios are poor; throughput stalls well below hardware limits.

Fix & prevent

Raise linger.ms (5–20 ms) and batch.size; compress with zstd or lz4; reduce partitions if records are spread too thin.

Explain it without notes

01

Why do batching and compression reinforce each other?

Practice

01

Run the compression comparison with your own JSON events and pick a codec.

Trade-offs

  • ↔

    Larger batches and linger raise throughput and compression but add milliseconds of latency.

  • ↔

    Stronger compression saves network and disk but costs CPU on producers and consumers.

Done when you can

  • I can tune batch.size, linger.ms and compression and read the metrics that prove it.