Topic 16.7
Capstone: A Distributed E-Commerce Event Platform
In one line
Bring everything together: services with transactional databases and outboxes, CDC into Kafka, domain topics keyed by entity, an orchestrated order saga, payment, inventory, shipping and notification consumers with idempotency, retries and DLTs, Kafka Streams analytics, a data warehouse, and full observability and DR.
Think of it like this
Running a national retail chain. Every store records its own sales (service databases), the company's logistics network carries every record reliably (Kafka), each department acts on what it needs (consumers), managers see live dashboards (Streams analytics), and head office keeps the full archive (warehouse).
Key ideas
- 01
Order service:
OrderCreated,OrderCancelledvia outbox →commerce.orders(key orderId). - 02
Payment:
PaymentRequested,PaymentAuthorized,PaymentFailed,Refundedonpayments.events(key orderId or paymentId), idempotency keys with providers. - 03
Inventory:
InventoryReserved,InventoryRejected,InventoryReleased(key orderId; stock updates keyed by sku on a separate topic), inbox dedup. - 04
Shipping:
ShipmentCreated,ShipmentDispatched,ShipmentDelivered. Notification: email, SMS and push consumers with dedup and retry topics. - 05
Analytics: Kafka Streams jobs for revenue per minute, top products per hour (windowed), funnel conversion; sinks to the warehouse.
- 06
Reliability: retries with backoff, retry topics, DLTs per consumer, outbox and inbox everywhere, saga orchestrator with timeouts. Observability: consumer lag (time), producer errors, broker health, processing latency, DLT growth, throughput. DR: MM2 to a second region; idempotent consumers.
Code & diagrams
Interview problem
The problem
Present the whole platform
Present the e-commerce event platform to a staff-level panel. For each major decision, answer: why Kafka, why this topic, these partitions, this key, this replication factor, these producer and consumer settings, at-least-once or exactly-once, this retry model, DLT, schema, retention, monitoring and DR strategy. Then: what happens if a broker dies, a consumer dies, the DB is down, a tenant becomes hot, or a region disappears?
You're given
- 5K orders/sec peak, 100× on sale days
- Money must never be lost or doubled
- Dashboards within 1 minute
- Two regions
Explain it without notes
Walk through what happens to an order when the inventory consumer crashes mid-reservation.
Practice
Build a vertical slice: order outbox → Kafka → payment and inventory consumers with inbox and DLT → Streams revenue-per-minute → verify with chaos tests.
Trade-offs
- ↔
The platform trades simplicity for decoupling, resilience and replayability; each pattern adds a moving part that must be monitored.
Done when you can
I can design and defend a complete Kafka-based e-commerce platform, including every failure scenario.