Command Palette

Search for a command to run...

Hectal
Phase 11Intermediate12 of 18 in Database Design

NoSQL and Specialised Databases

Choosing key-value, document, wide-column, graph and search stores; MongoDB document modeling; Cassandra query-first design; DynamoDB single-table design; Elasticsearch; Redis as a database; time-series and OLAP warehouses.

Each specialised store is excellent at a narrow set of access patterns and awkward outside it. This phase teaches the data model and the failure modes of each, so the choice follows from requirements rather than fashion.

0/7 · 0%
7 topics ~60 min 9 code blocks & diagrams
Start with the first topic
1
11.1

Choosing a Data Store: Key-Value, Document, Wide-Column, Graph, Search

Key-value stores answer get/put by key at massive scale; document stores keep whole aggregates together with flexible schemas; wide-column stores serve huge write-heavy datasets by partition key; graph databases traverse relationships; search engines rank free text. Start relational and add a specialised store when a measured access pattern needs it, keeping one source of truth.

8 min 1 diagram practice

2
11.2

MongoDB: Embedding, Referencing and Document Patterns

MongoDB stores BSON documents in collections. Model around what's read together: embed bounded, owned data (order lines in an order), reference unbounded or shared data (reviews, the product catalogue). Patterns such as bucket and extended reference handle growth and joins; compound, multikey and TTL indexes follow the same ESR rule as relational indexes; sharding needs a well-chosen shard key.

9 min 1 code practice

3
11.3

Cassandra: Query-First Modeling, Consistency Levels and Tombstones

Cassandra distributes rows by partition key over a token ring, replicates each partition to RF nodes, and orders rows inside a partition by clustering columns. You design one table per query, denormalising freely, keeping partitions bounded (~100 MB, ~100K rows is a common guide). Tunable consistency (ONE, QUORUM, LOCAL_QUORUM) sets the latency–consistency trade per request; deletes create tombstones that can slow reads.

9 min 1 code practice

4
11.4

DynamoDB: Keys, GSIs and Single-Table Design

DynamoDB stores items by partition key (and optional sort key) across auto-managed partitions, each with a throughput limit (about 3,000 read units and 1,000 write units per second). Model from access patterns: composite sort keys and overloaded GSIs serve many entity types in one table. Conditional writes provide uniqueness and optimistic locking; TTL expires items; transactions cover up to 100 items.

9 min 2 code practice

5
11.5

Elasticsearch: Inverted Indexes, Relevance and Reindexing

Elasticsearch (and OpenSearch) stores JSON documents in shards of Lucene segments with inverted indexes: analyzers turn text into terms, BM25 scores relevance, and filters narrow results cheaply. It's near-real-time (documents become searchable after a refresh, 1 s by default) and should be fed from the database as a derived index that can be rebuilt with an alias switch.

9 min 1 code practice

6
11.6

Redis as a Database: Structures, Persistence and Limits

Redis is an in-memory data-structure server: strings, hashes, lists, sets, sorted sets, streams, bitmaps and HyperLogLog, each with atomic commands. Persistence (RDB snapshots, AOF) and replication make it durable enough for many uses, but memory cost, asynchronous replication and eviction mean it's usually a cache or a store for ephemeral and derived data, not the only copy of critical records.

8 min 1 code practice

7
11.7

Time-Series Databases and OLAP Warehouses

Time-series data (metrics, sensor readings, events) is append-heavy, queried by time range and tags, and downsampled as it ages: use TimescaleDB, InfluxDB, Prometheus or ClickHouse. Analytical data belongs in a columnar warehouse (BigQuery, Snowflake, Redshift, ClickHouse) modeled as star schemas with fact and dimension tables, loaded by ETL or ELT, with slowly changing dimensions preserving history.

8 min 1 diagram 1 code practice