Topic 16.6
Log Aggregation and a 1,000-Database CDC Platform
In one line
Log aggregation uses lightweight collectors to batch and compress logs into Kafka, which buffers bursts before indexing and archival. A CDC platform runs Debezium connectors for many databases into Kafka, with schema governance, snapshots, deletes, and fan-out to search, the lake and analytics.
Think of it like this
A city's recycling system. Bins on every street (collectors) feed trucks (Kafka) that deliver to sorting plants (processors) and landfills or warehouses (storage). If a plant slows down, trucks wait in the yard instead of overflowing the streets.
Key ideas
- 01
Logs: collectors (Fluent Bit, Vector, OpenTelemetry Collector) batch and compress, send to Kafka with acks=1 or all depending on value; topics per environment or log class; keys null (spread) unless per-service ordering matters.
- 02
Log processing: parse, enrich, redact PII, route: hot logs to OpenSearch for search (days), everything to object storage (months), metrics derived in-stream. Kafka absorbs bursts during incidents, exactly when log volume spikes.
- 03
Hot partitions in logs: a chatty service with a service-name key overwhelms one partition; use null keys or service+random bucket.
- 04
CDC platform: Debezium connectors (one per database or group) on a large Connect cluster, topics per table with naming conventions, Avro with Schema Registry, snapshot orchestration (incremental snapshots to avoid long locks), heartbeat events to keep slots moving, and tombstones for deletes.
- 05
Governance: a catalogue of CDC topics with owners, PII classification and masking SMTs, schema change process with source teams, and backfill procedures for new consumers.
Code & diagrams
Interview problem
The problem
CDC platform for 1,000 databases
Design a CDC platform: 1,000 databases → CDC → Kafka → search, data lake and analytics. Discuss ordering, schema evolution, snapshots, deletes, replay and backfill.
Explain it without notes
Why does Kafka help log pipelines during incidents?
Practice
Design topic naming, partitioning and retention for logs from 500 services at 2 GB/s.
Trade-offs
- ↔
Kafka as a log buffer adds cost and a hop, but prevents log loss during bursts and decouples collection from indexing.
Done when you can
I can design log aggregation and multi-database CDC platforms on Kafka.