Command Palette

Search for a command to run...

Hectal
Phase 16Advanced17 of 18 in Apache Kafka

System Design with Kafka

Kafka vs Redis, RabbitMQ, SQS and Pub/Sub; anti-patterns and when not to use Kafka; FAANG-style designs for analytics pipelines, ride-hailing, activity feeds, payments, log aggregation and CDC platforms; and the e-commerce capstone.

Every design here follows the reasoning chain: business requirement → event characteristics → throughput and latency → durability → ordering → key → partitions → replication → producer and consumer semantics → retries and DLT → schemas → failure handling → observability → capacity → DR → trade-offs.

The goal is to decide precisely whether Kafka belongs, what it carries, and how it behaves when things fail.

0/7 · 0%
7 topics ~56 min 7 code blocks & diagrams
Start with the first topic
1
16.1

Kafka vs Redis, RabbitMQ, SQS, and Pub/Sub

Choose messaging technology from requirements, not habit: Kafka for durable, replayable, high-throughput, ordered-per-key streams with many consumers; RabbitMQ for flexible routing and per-message work queues; SQS and Google Pub/Sub for fully managed queues and fan-out; Redis for sub-millisecond state and lightweight streams.

8 min 1 code practice

2
16.2

Kafka Anti-Patterns and When Not to Use Kafka

Most Kafka pain comes from a known list: synchronous request/response over Kafka, huge payloads, treating Kafka as a database or cache, partition and topic explosions, ignored schemas, lag, idempotency and retention costs, random keys where order matters, one partition for huge loads, and every service consuming every event.

8 min 1 code practice

3
16.3

Designing Analytics Pipelines: Netflix Events and YouTube Analytics

High-volume analytics pipelines ingest client events through a gateway into Kafka, aggregate them with stream processors in windows, and land raw and aggregated data in a lake or warehouse. Key decisions: partition keys for aggregation, compression, retention, backpressure handling, and idempotent or exactly-once aggregation.

8 min 1 diagram practice

4
16.4

Ride-Hailing Trip Events and an Activity Feed

For ride-hailing, Kafka carries trip lifecycle events (keyed by trip), driver location streams (keyed by driver, short retention) and analytics, while geo matching lives in an in-memory geo index. For activity feeds, Kafka carries post, comment, like and connection events into fan-out and ranking processors that build per-user feed read models.

8 min 1 diagram practice

5
16.5

Designing a Payment Event Platform

Payment events (Created, Authorized, Captured, Failed, Refunded) demand strict per-payment ordering, no loss, no double effects, full auditability and long retention. Combine outbox publishing, RF 3 / acks=all / min.insync=2, idempotent consumers keyed on payment IDs, schemas with strict compatibility, DLTs with human review, and replayable history.

8 min 1 diagram practice

6
16.6

Log Aggregation and a 1,000-Database CDC Platform

Log aggregation uses lightweight collectors to batch and compress logs into Kafka, which buffers bursts before indexing and archival. A CDC platform runs Debezium connectors for many databases into Kafka, with schema governance, snapshots, deletes, and fan-out to search, the lake and analytics.

8 min 1 diagram practice

7
16.7

Capstone: A Distributed E-Commerce Event Platform

Bring everything together: services with transactional databases and outboxes, CDC into Kafka, domain topics keyed by entity, an orchestrated order saga, payment, inventory, shipping and notification consumers with idempotency, retries and DLTs, Kafka Streams analytics, a data warehouse, and full observability and DR.

8 min 1 diagram practice