Topic 11.2
MongoDB: Embedding, Referencing and Document Patterns
In one line
MongoDB stores BSON documents in collections. Model around what's read together: embed bounded, owned data (order lines in an order), reference unbounded or shared data (reviews, the product catalogue). Patterns such as bucket and extended reference handle growth and joins; compound, multikey and TTL indexes follow the same ESR rule as relational indexes; sharding needs a well-chosen shard key.
Think of it like this
A patient file. The current prescription and allergies are stapled into the folder (embedded), while X-ray archives live in the radiology department with a reference number (referenced), because they're huge and shared.
Key ideas
- 01
Embed when data is owned by the parent, bounded in size, and read together (addresses in a user, lines in an order). Reference when it's unbounded (comments on a viral post), shared by many parents, or changes independently. Documents are capped at 16 MB, and large documents rewrite and transfer poorly.
- 02
Extended reference: store the referenced entity's ID plus the few fields you display (customer name on an order) to avoid
$lookup; accept that the copy may be stale or is "as of order time". - 03
Bucket pattern: group many small items into one document per time window (one document per sensor per hour with an array of readings), which cuts index size and document count. MongoDB time-series collections do this automatically.
- 04
Indexes: single-field, compound (ESR: equality, sort, range), multikey (one entry per array element), text, TTL (
expireAfterSecondsfor automatic deletion), partial and wildcard.explain("executionStats")showstotalKeysExaminedvsnReturned. - 05
Replica sets: one primary, secondaries replicate the oplog; write concern
majorityand read concernmajorityfor durability. Sharding: chunks (ranges of the shard key) are balanced across shards; a monotonically increasing shard key creates a hot shard (use hashed or compound keys).
Code & diagrams
// embed owned, bounded data; extended reference to the customer
db.orders.insertOne({
_id: ObjectId(),
customer: { _id: ObjectId("66f1a2b3c4d5e6f708192a3b"), name: "Asha Rao" },
status: "placed",
createdAt: ISODate("2026-09-14T10:15:00Z"),
lines: [
{ sku: "TSHIRT-M-RED", qty: 2, unitPrice: NumberDecimal("499.00") },
{ sku: "CAP-BLUE", qty: 1, unitPrice: NumberDecimal("299.00") }
],
total: NumberDecimal("1297.00")
});
// ESR compound index for "a customer's orders by status, newest first"
db.orders.createIndex({ "customer._id": 1, status: 1, createdAt: -1 });
// bucket pattern: one document per sensor per hour
db.readings.updateOne(
{ sensorId: "s-17", hour: ISODate("2026-09-14T10:00:00Z"), count: { $lt: 3600 } },
{ $push: { r: { t: ISODate("2026-09-14T10:15:02Z"), v: 21.7 } }, $inc: { count: 1 } },
{ upsert: true }
);
// TTL index: delete sessions 24h after lastSeen
db.sessions.createIndex({ lastSeen: 1 }, { expireAfterSeconds: 86400 });Interview problem
The problem
Model a blog with viral posts in MongoDB
Posts have an author, tags and comments. Most posts get a few comments; viral posts get 500K. The post page shows the post, the author's name and the newest 20 comments, with pagination. Design the collections.
When it breaks
Unbounded embedded arrays
What you see
A viral post's embedded comments hit the 16 MB document limit; before that, every update rewrites a multi-MB document and replication traffic explodes.
Fix & prevent
Reference unbounded children in their own collection; embed only a bounded subset (subset pattern).
Explain it without notes
What decides embedding vs referencing?
Practice
Choose a shard key for an orders collection queried mostly by customer.
Trade-offs
- ↔
Embedding gives single-read pages and atomic updates at the cost of duplication and document growth; referencing is normalized but needs more queries.
Done when you can
I can model MongoDB documents with embedding, references, bucket and extended reference patterns and choose indexes and shard keys.