Command Palette

Search for a command to run...

Hectal
PHASE 12Intermediate ~8 min· topic 3 of 4

Topic 12.3

Polyglot Persistence: One Truth, Many Copies

In one line

Real systems combine PostgreSQL, Redis, Kafka, Elasticsearch, object storage and sometimes Cassandra or MongoDB. For each piece of data, name the source of truth and the synchronisation mechanism (same transaction, outbox/CDC, cache-aside, batch) for every copy, with its consistency, durability and staleness budget.

0/4 · 0%

Think of it like this

A company's official records are in the registry office (source of truth); departments keep photocopies for convenience. Everyone knows which copy wins in a dispute, and there's a process for updating photocopies.

Key ideas

  1. 01

    Typical roles: PostgreSQL = transactional source of truth; Redis = cache, sessions, counters, rate limits; Kafka = durable change log between systems; Elasticsearch = search projection; object storage = blobs, archives and the data lake; warehouse = analytics.

  2. 02

    For each component record: source of truth or derived; durability (what's lost on node failure); consistency with the source (sync, seconds, minutes); query capability; latency; scale limits; cost; synchronisation mechanism and how to rebuild it.

  3. 03

    Synchronisation: CDC or outbox for projections (search, warehouse, caches); cache-aside with invalidation for caches; application-level sagas for cross-service data; never unguarded dual writes.

  4. 04

    Every derived store needs a rebuild path (replay the Kafka topic from the beginning or retention, or snapshot + stream) and a reconciliation check (counts, checksums, sampled comparisons).

  5. 05

    Each extra store costs operations, on-call knowledge, security review, backups and failure modes. Add one only when an access pattern demands it.

Code & diagrams

data-inventory.mdmarkdown
| Data            | Source of truth | Copies                  | Sync            | Staleness | Rebuild           |
|-----------------|-----------------|-------------------------|-----------------|-----------|-------------------|
| Orders          | PostgreSQL      | Kafka, warehouse        | outbox + CDC    | < 5 s / 1 h | replay / re-snapshot |
| Product catalog | PostgreSQL      | Redis, Elasticsearch    | CDC -> Kafka    | < 5 s     | reindex + alias   |
| Sessions        | Redis           | none                    | -               | -         | users log in again |
| Images          | S3              | CDN                     | cache headers   | minutes   | purge CDN         |
| Clickstream     | Kafka -> S3     | ClickHouse              | stream job      | < 1 min   | replay from S3    |

Interview problem

The problem

Design data architecture for a marketplace

A marketplace has listings, orders, payments, messaging between buyers and sellers, search, recommendations and analytics. Propose the stores, the source of truth for each dataset, and how they synchronise.

When it breaks

Two services both think they own customer data

What you see

Profile edits in one service are overwritten by nightly syncs from the other; support can't say which address is correct.

Fix & prevent

One owning service per entity; others consume its events and treat their copies as read-only projections.

Explain it without notes

01

Why must every derived store be rebuildable?

Practice

01

Fill the data-inventory table for a food delivery app with at least five datasets.

Trade-offs

  • ↔

    Right tool per access pattern versus operational complexity and consistency work across stores.

Done when you can

  • I can name the source of truth and the sync mechanism for every copy in an architecture.