Command Palette

Search for a command to run...

Hectal
Phase 1Beginner2 of 18 in Apache Kafka

Topics, Partitions & Keys

How topics are configured and named, why partitions exist, how keys decide ordering and placement, and how to reason about partition count instead of guessing a number.

Partitions are the most important idea in Kafka. They give you parallelism and scale, and they define exactly where ordering holds. Keys decide which partition each record goes to, which means your key choice is really an ordering and load-balancing decision.

Get these right and most of Kafka falls into place; get them wrong and no tuning will save you, because partition counts are hard to change later without breaking key ordering.

0/5 · 0%
5 topics ~41 min 6 code blocks & diagrams
Start with the first topic
1
1.1

Topics and Topic Configuration

A topic is a named, partitioned log with its own configuration: partition count, replication factor, retention, cleanup policy, min.insync.replicas, segment size and more. Topic-level settings override broker defaults, and a few of them decide durability and cost more than anything else.

7 min 1 code practice

2
1.2

Topic Naming and Organisation

Topic names are a shared API: choose a convention that encodes domain, entity or event, and version, keep environments in separate clusters rather than name prefixes where possible, and decide between one topic per entity stream and one topic per event type based on ordering and consumer needs.

7 min 1 code practice

3
1.3

Partitions and Ordering

A partition is one ordered, append-only log; Kafka guarantees order only within a partition. Records with the same key go to the same partition, so keying by entity ID (for example orderId) keeps each entity's events in order while different entities are processed in parallel.

9 min 1 diagram 1 code practice

4
1.4

Choosing the Partition Count

Partition count sets maximum consumer parallelism and spreads load, but each partition costs broker memory, file handles, replication work, metadata, failover and rebalance time. Size it from target throughput divided by measured per-partition throughput for producers and consumers, plus headroom, and plan for growth because increasing it later remaps keys.

9 min 1 code practice

5
1.5

Partitioners and Partition Key Design

The key decides three things at once: which partition a record lands in, the scope of ordering, and how evenly load spreads. Choose the narrowest entity whose events must be ordered (order, account, device), check the key distribution for skew, and use custom partitioning only when you understand what you're giving up.

9 min 1 code practice