Command Palette

Search for a command to run...

Hectal
PHASE 0Beginner ~9 min· topic 1 of 5

Topic 0.1

What Kafka Is: A Distributed Commit Log

In one line

Kafka stores streams of records in topics, split into partitions, each an append-only, ordered, replicated log on disk. Producers append; consumers read by offset and track their own position. Because records aren't deleted on read, many consumers can process the same data independently and replay history.

0/5 · 0%

Think of it like this

A bank's transaction ledger. Every transaction is written on the next line and never erased. The fraud team, the statements team and the auditors each read the same ledger at their own speed with their own bookmark. Nobody "takes" a line away from the others, and an auditor can go back and re-read last month.

Key ideas

  1. 01

    Traditional queue (RabbitMQ, SQS): a message is delivered to one consumer and removed once acknowledged. Commit log (Kafka): records stay for a retention period (hours, days, forever); each consumer group keeps its own offset. Event streaming platform: the log plus the ecosystem around it: Connect for integration, Streams for processing, Schema Registry for contracts.

  2. 02

    Roles Kafka plays: event bus between microservices, durable buffer that absorbs spikes, integration layer (databases → Kafka → search, warehouse, cache), change data capture transport, analytics pipeline, audit or event store, and the backbone for stream processing.

  3. 03

    What makes it scale: a topic is split into partitions spread across brokers, so writes and reads for different partitions happen in parallel on different machines. Ordering is guaranteed within a partition, not across the topic.

  4. 04

    What makes it durable: each partition is replicated to several brokers; a write can be acknowledged only after all in-sync replicas have it (acks=all), and data is kept on disk according to retention settings.

  5. 05

    What Kafka isn't: a database you query by field, a request/response RPC system, a task queue with per-message priorities and delays (though Kafka 4.x adds early "share groups" for queue-like consumption), or a cache. Using it as those leads to pain.

Code & diagrams

kafka-log.mermaiddiagram
Rendering diagram…
queue-vs-log.txttext
                     Traditional queue          Kafka (commit log)
After consumption    message removed            record stays until retention
Consumers            compete for messages       each group reads everything
Replay               no                         yes, reset the offset
Ordering             per queue (often loose)    per partition, strict
Scaling consumers    add workers                add partitions + consumers
Per-message ack      yes                        offset commit (position)
Typical retention    until consumed             hours to forever
Throughput           thousands-100Ks msg/s      millions msg/s per cluster

Interview problem

The problem

Queue, log, or streaming platform?

Three teams ask for "a queue": (1) image resizing jobs, (2) order events that email, analytics, fraud and search must all process, (3) a real-time fraud score computed from the last 10 minutes of card swipes. Which need Kafka, and why?

The interviewer follows up

01

Can Kafka replace a database?

Explain it without notes

01

Explain the difference between a queue and a commit log in one example.

Practice

01

List three systems at your company that exchange data. For each, would a queue or a log fit better, and why?

Trade-offs

  • ↔

    Kafka gives durable, replayable, high-throughput streams, at the cost of operating a distributed system and designing partitions and keys carefully.

Done when you can

  • I can describe Kafka as a partitioned, replicated commit log.

  • I can explain when a log beats a queue and when it doesn't.