Command Palette

Search for a command to run...

Hectal
PHASE 2Intermediate ~10 min· topic 4 of 4

Topic 2.4

Memory Optimisation and Capacity Planning

In one line

Estimate Redis memory before you build, then shrink it with shorter keys, bucketed hashes, compact encodings, compression and TTLs, and size clusters with headroom for replication, fork and growth.

0/4 · 0%

Think of it like this

Packing a moving truck. You measure the furniture first, disassemble what you can, pack small things into boxes instead of loose, and rent a truck with some spare room, because the sofa never fits the way you thought it would.

Key ideas

  1. 01

    Estimate per item: key bytes + value bytes + overhead (~50–90 bytes per top-level key, ~10–30 bytes per element inside a listpack, ~60–100 bytes per element in a large hash, set or sorted set). Then multiply by count, add ~20–30% for fragmentation and buffers, and verify empirically: load 100K realistic items and measure used_memory growth.

  2. 02

    Reduction techniques, from easiest: shorter key prefixes, TTLs on everything temporary, removing unused fields, storing numbers as integers (the int encoding), bucketing many small keys into hashes kept under listpack thresholds, compact binary serialisation (MessagePack, Protobuf) instead of verbose JSON, and compressing large values (LZ4, Snappy, zstd) above ~1 KB.

  3. 03

    The classic bucketing example (Instagram, 2011): mapping 300M media IDs to user IDs as plain keys needed ~21 GB; storing them in hashes of 1,000 fields each (mediabucket:{id/1000}, field {id}) with listpack-style encoding brought it to under 5 GB. The trade-off is that per-field TTL (before 7.4) and key-level operations no longer apply to each item.

  4. 04

    Plan capacity per shard: keep each primary's dataset in a range that makes failover and resync quick (commonly under ~25 GB per shard), and leave headroom: dataset × 1.3 (fragmentation/buffers) plus fork copy-on-write, so machine RAM ≈ 1.5–2× dataset for write-heavy instances with persistence.

  5. 05

    Throughput is part of capacity too: one shard does roughly 100K–300K simple ops/sec without pipelining. Size shard count from both memory and ops/sec, and add replicas for read scaling and failover (they double memory cost).

  6. 06

    Cost matters: cloud RAM is expensive. Compare the cost of caching everything to the cost of the misses (database load, latency). Often caching the hottest 10–20% gives most of the benefit.

Code & diagrams

estimate.pypython

Measure, don't guess: load a realistic sample and extrapolate.

import json, random, redis

r = redis.Redis()
r.flushdb()
base = r.info("memory")["used_memory"]

N = 100_000
pipe = r.pipeline(transaction=False)
for i in range(N):
    session = {"userId": random.randint(1, 10**7), "role": "user",
               "device": "Android 14", "ip": "10.0.0.1", "loginTime": 1727520000}
    pipe.hset(f"session:{i:016x}", mapping=session)
    pipe.expire(f"session:{i:016x}", 1800)
    if i % 1000 == 0:
        pipe.execute()
pipe.execute()

per_item = (r.info("memory")["used_memory"] - base) / N
print(f"{per_item:.0f} bytes per session")
print(f"100M sessions ~ {per_item * 100e6 / 2**30:.1f} GiB before headroom")
# 208 bytes per session
# 100M sessions ~ 19.4 GiB before headroom
bucketing.redisredis
# Before: one key per item (~90 bytes overhead each)
SET media:1155315 939
SET media:1155316 3022

# After: 1,000 items per hash, kept under hash-max-listpack-entries
CONFIG SET hash-max-listpack-entries 1024
HSET mediabucket:1155 1155315 939
HSET mediabucket:1155 1155316 3022
HGET mediabucket:1155 1155315
"939"

Interview problem

The problem

Size a Redis cluster for a product cache

You're caching 40M product documents (average 2 KB of JSON each) and serving 400K reads/sec and 20K writes/sec at peak. Estimate memory and shard count, propose reductions, and describe headroom.

You're given

  • 40M products
  • 2 KB JSON each
  • 400K reads/sec
  • 20K writes/sec
  • One replica per shard

The interviewer follows up

01

What if the business insists on caching all 40M products?

When it breaks

Estimating memory from value sizes only

What you see

The cluster fills at 60% of the planned item count because per-key overhead, fragmentation and replication buffers weren't counted.

Fix & prevent

Load a realistic sample and measure used_memory per item; add 30% for fragmentation and buffers.

One huge shard (100 GB) instead of several smaller ones

What you see

A failover triggers a full resync of 100 GB, which takes many minutes and saturates the network, and forks take seconds.

Fix & prevent

Keep per-shard datasets moderate (tens of GB at most) and scale out with more shards.

Explain it without notes

01

How would you estimate Redis memory for a new feature before building it?

02

Explain the hash-bucketing technique and its downsides.

Practice

01

Run the estimate.py script in your lab and extrapolate to 100M sessions. How many shards would you use?

Trade-offs

  • ↔

    Compression saves memory and network bandwidth but costs CPU on the clients and makes values opaque to Redis commands.

  • ↔

    Caching everything maximises hit ratio; caching the hot set costs far less, with slightly more misses.

Done when you can

  • I can estimate memory for a dataset and validate it empirically.

  • I can list at least five memory reduction techniques.

  • I can size shards for both memory and ops/sec with headroom.