Command Palette

Search for a command to run...

Hectal
PHASE 13Advanced ~10 min· topic 2 of 5

Topic 13.2

Big Keys: Finding and Fixing Them Safely

In one line

A big key (a huge string or a collection with millions of elements) makes every whole-key operation slow, blocks the server, inflates replication and migration, and can stall Redis for seconds when deleted. Find them with --bigkeys/--memkeys, read them with SCAN-family commands, delete them with UNLINK, and split them for good.

0/5 · 0%

Think of it like this

One enormous filing cabinet that takes three people and ten minutes to move. Every time someone needs one folder from it, they have to wheel the whole cabinet out. Throwing it away also blocks the corridor for an hour.

Key ideas

  1. 01

    What counts as big: strings over ~1 MB, or collections over ~10,000 elements, as rough alarms; anything where a whole-key command takes more than a millisecond or two.

  2. 02

    Why they hurt: HGETALL, SMEMBERS, LRANGE 0 -1, ZRANGE 0 -1 are O(N) and build giant replies; DEL is O(N) for collections (freeing each element); MIGRATE during resharding serialises the whole key; replication and AOF rewrite spike; a failover or restart takes longer.

  3. 03

    Finding them: redis-cli --bigkeys (largest by element count per type), --memkeys (largest by memory), MEMORY USAGE key, RDB analysis tools offline, and SLOWLOG entries showing O(N) commands on the same key.

  4. 04

    Reading safely: HSCAN, SSCAN, ZSCAN with COUNT iterate in small steps; range commands with bounded ranges; HMGET for specific fields.

  5. 05

    Deleting safely: UNLINK key removes the key from the keyspace immediately and frees memory in a background thread. lazyfree-lazy-user-del yes makes DEL behave like UNLINK; lazyfree-lazy-eviction, lazyfree-lazy-expire and lazyfree-lazy-server-del do the same for eviction, expiry and internal deletes (some newer versions, such as Valkey 8, enable these by default). Alternatively trim a collection in chunks (HSCAN + HDEL batches) before deleting.

  6. 06

    Fixing permanently: split by bucket (user:123:events:{0..63} or by time: events:2026-09-28), cap sizes (LTRIM, XADD MAXLEN, ZREMRANGEBYRANK), compress large strings, or move large blobs to object storage and keep only references in Redis.

Code & diagrams

bigkeys.shbash
redis-cli --bigkeys -i 0.01          # -i sleeps between SCAN batches to limit impact
[00.00%] Biggest hash   found so far '"huge:user"' with 4812331 fields
[00.00%] Biggest string found so far '"report:2026"' with 48203112 bytes
-------- summary -------
Biggest   hash found '"huge:user"' has 4812331 fields
Biggest string found '"report:2026"' has 48203112 bytes

redis-cli --memkeys
Biggest hash found '"huge:user"' has 612.4 MB

redis-cli MEMORY USAGE huge:user SAMPLES 5
(integer) 642211840
redis-cli SLOWLOG GET 1
1) 1) (integer) 991
   2) (integer) 1727520000
   3) (integer) 2310442              # 2.3 seconds
   4) 1) "HGETALL"
      2) "huge:user"
split-and-delete.pypython
def migrate_big_hash(r, src="huge:user", buckets=256):
    cursor = 0
    while True:
        cursor, fields = r.hscan(src, cursor, count=500)
        if fields:
            pipe = r.pipeline(transaction=False)
            for f, v in fields.items():
                b = hash(f) % buckets              # use a stable hash (e.g. crc32) in real code
                pipe.hset(f"user:events:{{u123}}:{b}", f, v)
            pipe.execute()
        if cursor == 0:
            break
    r.unlink(src)                                  # frees memory in a background thread

Interview problem

The problem

HGETALL huge:user takes seconds

You discover HGETALL huge:user takes several seconds and causes latency spikes for everyone. Fix it without bringing down production.

You're given

  • ~5M fields, ~600 MB
  • Called by a user-profile endpoint
  • Can't take downtime

The interviewer follows up

01

One Redis key contains 50 GB. What do you do?

When it breaks

DEL on a 10M-element sorted set

What you see

The main thread spends seconds freeing elements; every client times out; Sentinel or Cluster may start a failover mid-delete.

Fix & prevent

UNLINK (or enable lazyfree-lazy-user-del); for expired or evicted big keys, enable the other lazyfree options.

Explain it without notes

01

Why is UNLINK safer than DEL for large keys?

Practice

01

Create a 2M-member set in the lab, time DEL vs UNLINK on copies of it with redis-cli --latency running in another terminal.

Trade-offs

  • ↔

    Splitting big keys spreads load and makes operations cheap, at the cost of more keys and multi-key reads.

Done when you can

  • I can find big keys by elements and by memory.

  • I can read, migrate and delete big keys without blocking Redis.