Topic 13.3
Performance Engineering: Complexity, Round Trips, Connections
In one line
Redis performance comes down to command complexity (avoid unbounded O(N)), round trips (batch, pipeline, script), payload size, connection management, CPU saturation of the main thread, and memory pressure. SCAN-family commands replace dangerous full-collection reads.
Think of it like this
A fast cashier's speed depends on how many items each customer brings (command complexity), how often customers walk away and come back (round trips), how big the bags are (payload), and how many people crowd the till at once (connections).
Key ideas
- 01
Complexity ladder: O(1) (
GET,HGET,SADD), O(log N) (ZADD,ZRANK), O(log N + M) ranges with small M, and O(N) (KEYS,SMEMBERS,HGETALL,LRANGE 0 -1,SUNION,FLUSHALLwithout ASYNC). O(N) on small N is fine; on unbounded N it's an outage waiting to happen. - 02
Dangerous commands on big data:
KEYS *,SMEMBERS huge-set,HGETALL huge-hash,LRANGE 0 -1,ZRANGE 0 -1,DELon huge collections,FLUSHALL/FLUSHDBwithoutASYNC,DEBUG,MONITOR. Block risky ones for application users with ACLs (-@dangerous,-keys). - 03
SCAN semantics:
SCAN cursor MATCH p COUNT nreturns a new cursor and some keys; iteration ends when the cursor returns to 0. Guarantees: every key present for the whole iteration is returned at least once; keys may be returned more than once; keys added or removed during the scan may or may not appear.COUNTis a hint about work per call, not an exact number.MATCHfilters after fetching, so sparse patterns may return empty pages. - 04
Round trips and payloads: batch with
MGET/HMGET, pipelines and Lua; keep values small (compress over ~1 KB); a 1 MB value costs ~1 ms of network time on a 10 Gbps link, which adds up at thousands of requests/sec. - 05
CPU: the main thread saturates at 100% of one core. Symptoms: rising latency across all commands. Causes: expensive commands, Lua scripts, too many small commands (use pipelining), TLS overhead (consider I/O threads), or hot keys. Check
INFO commandstatsforusec_per_calland call counts, andINFO cpu. - 06
Connections: every connection costs memory; connection storms (thousands of new connections per second, for example from short-lived serverless functions or no pooling) burn CPU on accept and TLS handshakes. Reuse connections, pool properly, and watch
connected_clientsandtotal_connections_received.
Code & diagrams
127.0.0.1:6379> INFO commandstats
cmdstat_get:calls=982112233,usec=781231221,usec_per_call=0.80,rejected_calls=0,failed_calls=0
cmdstat_hgetall:calls=221023,usec=412222011,usec_per_call=1865.10,... # <-- expensive
cmdstat_evalsha:calls=8821021,usec=91022113,usec_per_call=10.32,...
127.0.0.1:6379> SCAN 0 MATCH session:* COUNT 1000
1) "38712"
2) 1) "session:9f3c"
2) "session:a1b2"
...
127.0.0.1:6379> ACL SETUSER app on >s3cret ~app:* +@all -@dangerous -keys
OKInterview problem
The problem
Redis CPU at 100%
Redis CPU is pinned at 100% on one core and latency is climbing. How do you debug it?
The interviewer follows up
Would enabling I/O threads help?
Explain it without notes
What exactly does SCAN guarantee and not guarantee?
Practice
Read INFO commandstats on a real (or lab) instance and list the top 3 commands by total CPU time.
Trade-offs
- ↔
Batching and scripting reduce round trips but make individual operations larger; keep each unit of work bounded.
Done when you can
I know the complexity of the commands I use and which ones to ban on big data.
I can use SCAN correctly and debug CPU saturation with commandstats.