Topic 15.3
Event Sourcing
In one line
Event sourcing stores state as an append-only sequence of domain events; current state is the result of replaying them, often from a snapshot. Kafka can distribute and retain events, but an event-sourced system also needs per-aggregate reads, optimistic concurrency and snapshots, which a dedicated event store or database usually provides better.
Think of it like this
A bank statement. Your balance isn't stored as a single number that gets overwritten; it's derived from every deposit and withdrawal since the account opened. You can explain any balance by replaying the history.
Key ideas
- 01
Model:
AccountCreated,MoneyDeposited,MoneyWithdrawnfor aggregateaccount-123. State = fold(events). New commands are validated against the current state, then produce new events. - 02
Benefits: full audit trail, time travel (state at any point), new projections from history, and natural fit with event-driven integration.
- 03
Needs: append events for one aggregate with optimistic concurrency ("expected version 17"), read all events for one aggregate quickly, snapshots to avoid replaying millions of events, and immutable events with upcasting for schema evolution.
- 04
Kafka's fit: great for publishing events to other services and retaining them (compacted snapshots, infinite retention with tiered storage), but reading one aggregate's history requires scanning a partition, and Kafka has no "append if version = N" check. Common designs use a database or event store (EventStoreDB, PostgreSQL events table) as the source of truth and Kafka to distribute events (via outbox/CDC).
- 05
Source-of-truth decision: if Kafka is the store, retention must be infinite (or snapshot + compaction), backups and DR must be treated like a database, and GDPR erasure needs crypto-shredding (encrypt personal data per user and delete the key).
Code & diagrams
public final class Account {
private long balanceMinor;
private int version;
public static Account replay(Snapshot snap, List<AccountEvent> events) {
Account a = snap == null ? new Account() : snap.toAccount();
events.forEach(a::apply);
return a;
}
public List<AccountEvent> withdraw(long amountMinor) {
if (amountMinor > balanceMinor) throw new InsufficientFunds();
return List.of(new MoneyWithdrawn(amountMinor)); // validate, then emit
}
void apply(AccountEvent e) {
switch (e) {
case AccountCreated c -> balanceMinor = 0;
case MoneyDeposited d -> balanceMinor += d.amountMinor();
case MoneyWithdrawn w -> balanceMinor -= w.amountMinor();
}
version++;
}
}
// Append with optimistic concurrency (event store / DB):
// INSERT INTO events(aggregate_id, version, ...) VALUES ('account-123', 18, ...) -- unique (aggregate_id, version)Explain it without notes
Why is Kafka alone an awkward event store for aggregates?
Practice
Implement an event-sourced account with a snapshot every 100 events and measure load time with and without snapshots.
Trade-offs
- ↔
Event sourcing gives auditability and replay, at the cost of complexity, eventual consistency in read models and schema-evolution discipline forever.
Done when you can
I can explain event sourcing, snapshots, concurrency and when Kafka is or isn't the event store.