Command Palette

Search for a command to run...

Hectal
Phase 13Advanced14 of 18 in Apache Kafka

Production Operations

Security (TLS, SASL, ACLs, encryption), quotas and multi-tenancy, capacity planning with real arithmetic, large messages and the claim-check pattern, and disaster recovery across regions.

Running Kafka for many teams means securing it, stopping one tenant from starving others, sizing it with numbers rather than hope, keeping payloads sane, and having a tested answer for "the region is gone".

Security content is defensive: what to protect, how to configure it, and how to verify it.

0/5 · 0%
5 topics ~39 min 7 code blocks & diagrams
Start with the first topic
1
13.1

Security: Authentication, Authorization, Encryption

Secure Kafka in layers: TLS for encryption in transit, SASL (SCRAM, OAUTHBEARER, GSSAPI) or mutual TLS for authentication, ACLs with least privilege per principal on topics, groups and the cluster, encryption at rest for disks and backups, and secret and certificate rotation.

8 min 2 code practice

2
13.2

Quotas and Multi-Tenancy

Quotas cap produce bytes/sec, fetch bytes/sec and request CPU time per user or client ID; brokers throttle offenders by delaying responses instead of rejecting them. Combined with tenant-aware topic strategies, they keep noisy neighbours from degrading everyone.

7 min 1 code practice

3
13.3

Capacity Planning with Real Numbers

Size a cluster from throughput (MB/s in), replication factor, consumer fan-out, retention and compression: they give you disk, network and broker counts. Add headroom for failures (a lost broker's load moves to others), recovery traffic and growth, then validate with a load test.

8 min 1 code practice

4
13.4

Large Messages and the Claim-Check Pattern

Kafka's default max message size is about 1 MB. Very large records (20 MB JSON) cause memory pressure, slow replication, consumer instability and head-of-line blocking. Store large payloads in object storage and send a small event with a reference (claim-check), or split and compress.

7 min 1 diagram practice

5
13.5

Disaster Recovery and Multi-Region Kafka

Kafka clusters live in one region; for regional disasters you replicate topics to another cluster with MirrorMaker 2 (or Cluster Linking, MSK Replicator and similar). Plan RPO (replication lag), RTO (failover procedure), consumer offset translation, duplicate handling, and whether you run active-passive or active-active.

9 min 1 diagram 1 code practice