Command Palette

Search for a command to run...

Hectal
PHASE 16Advanced ~8 min· topic 2 of 7

Topic 16.2

Kafka Anti-Patterns and When Not to Use Kafka

In one line

Most Kafka pain comes from a known list: synchronous request/response over Kafka, huge payloads, treating Kafka as a database or cache, partition and topic explosions, ignored schemas, lag, idempotency and retention costs, random keys where order matters, one partition for huge loads, and every service consuming every event.

0/7 · 0%

Think of it like this

Using a freight train for a pizza delivery. It can technically carry the pizza, but everything about it is wrong for the job.

Key ideas

  1. 01

    Request/response over Kafka: a service publishes a request and blocks waiting for a reply event; you get REST's coupling plus queueing latency and correlation complexity. Use HTTP/gRPC for queries; use events for facts.

  2. 02

    Kafka as a database: querying by field, updating records, joins. Build read models (Topic 15.4) instead.

  3. 03

    Kafka as a cache: short-lived key lookups belong in Redis or a local cache.

  4. 04

    Partition and topic explosions: thousands of partitions per topic "for safety", or a topic per user/tenant/device, overload metadata, recovery and rebalances.

  5. 05

    Ignoring contracts and operations: no schemas, no consumer lag monitoring, no idempotency ("exactly-once will handle it"), no retention budget, secrets or personal data in payloads, random keys when ordering matters, a single partition for a massive stream, every microservice subscribing to everything.

  6. 06

    When not to use Kafka at all: low-volume point-to-point work queues (use SQS/RabbitMQ), synchronous queries, tiny teams without capacity to operate it (use a managed service or a simpler queue), strict per-message scheduling and priorities, and cases where a database table plus a poller is enough.

Code & diagrams

anti-patterns.txttext
Anti-pattern                              Why it hurts                           Instead
Request/response over Kafka               latency + coupling + correlation code  HTTP/gRPC for queries
20 MB payloads                            memory, replication, head-of-line      claim-check (object storage + event)
Kafka as a queryable DB                   no indexes or ad-hoc queries           read models (search, DB, cache)
1 topic per user / tenant                 millions of partitions                 shared topics keyed by id, quotas
5,000 partitions "to be safe"             slow recovery, small batches           size from measured throughput
No schemas                                silent consumer breakage               registry + compatibility in CI
Random keys, ordering required            events reordered                       key by entity id
"EOS means no duplicates anywhere"        double emails and charges              idempotent side effects
Unbounded retention by default            runaway storage cost                   retention per topic, tiered storage
Every service consumes every event        hidden coupling, wasted fetches        subscribe only to owned needs

Explain it without notes

01

Why is request/response over Kafka usually a bad idea?

Practice

01

Audit one Kafka-based system you know against the anti-pattern list and rank the top three issues.

Trade-offs

  • ↔

    Kafka is an excellent default for event streams and a poor one for everything else.

Done when you can

  • I can name Kafka anti-patterns and explain when not to use Kafka.