Topic 17.3
Question Bank: Scaling, Distributed Systems and Production
In one line
Eighteen questions on replication, sharding, hot partitions, CAP and PACELC, quorums, consistency, outbox and CDC, zero-downtime migrations, backups, RPO/RTO and monitoring, with model answers aimed at senior-level depth.
Think of it like this
The second round of a quiz, where questions are about judgement rather than definitions.
Key ideas
- 01
At senior level, answers are expected to include numbers, failure behaviour and trade-offs, not just definitions.
- 02
When unsure, reason from first principles: where is the data, who writes it, how does it get copied, and what happens when a copy is late or missing.
Code & diagrams
Definition -> mechanism -> when to use -> what breaks -> how you'd detect it -> alternativeExplain it without notes
Replication vs sharding?
How do you choose a shard key?
What causes hot partitions and how do you fix them?
How do you reshard without downtime?
How do you handle cross-shard queries?
Explain CAP.
Explain PACELC.
What is a quorum?
Strong vs eventual consistency?
Leader-based vs leaderless replication?
How do you guarantee read-your-writes with replicas?
Why do dual writes fail and what's the fix?
How do you perform a zero-downtime migration?
How do you recover a failed or damaged database?
What are RPO and RTO?
How do you monitor a database?
What happens when VACUUM can't keep up?
When would you pick Cassandra or DynamoDB over PostgreSQL?
Practice
Pick six questions at random and answer each with a number and a failure mode.
Trade-offs
- ↔
Depth beats breadth in senior interviews: fewer topics, explained with mechanisms, numbers and failure behaviour.
Done when you can
I can answer scaling, distributed and production questions with numbers and failure handling.