Topic 8.1
Client Architecture: Pools, Multiplexing, Timeouts, Retries
In one line
Redis clients either pool many connections (Jedis, redis-py) or multiplex many requests over one connection (Lettuce, ioredis, node-redis). Either way, you must set timeouts, bound retries, size pools deliberately, and handle topology changes, or a slow Redis becomes a full application outage.
Think of it like this
A taxi rank. A pool is a fixed number of taxis: when all are out, the next passenger waits. Multiplexing is a bus: everyone boards the same vehicle and gets off at their stop, which is efficient, until one passenger with a huge suitcase (a blocking command or a giant reply) holds everyone up.
Key ideas
- 01
Pooled clients: one connection serves one command at a time; concurrency needs a pool. Pool size ≈ peak concurrent Redis calls per instance, not "as large as possible": 500 instances × 100 connections = 50,000 connections, well above the default
maxclientsof 10,000, and each connection costs Redis memory. - 02
Multiplexed clients: many threads share one (or a few) connections; commands are pipelined automatically and replies matched in order. Very efficient, but blocking commands (
BLPOP,XREAD BLOCK,SUBSCRIBE) andMULTI/WATCHneed dedicated connections, and one huge reply delays every request on that connection. - 03
Timeouts: set a connect timeout (~1 s) and a command timeout sized to your latency budget (often 50–500 ms). Without them, a stalled Redis makes every request thread wait, exhausting thread pools across the app.
- 04
Retries: retry only idempotent operations (reads,
SET,DEL), with jittered exponential backoff and a small cap. RetryingINCRorLPUSHafter a timeout can apply it twice, because the first attempt may have succeeded. Use a circuit breaker so a dead Redis fails fast instead of piling up retries. - 05
Topology: Sentinel-aware clients ask Sentinel for the current primary; cluster-aware clients cache the slot map, follow
MOVED/ASK, and refresh topology periodically and on errors. Enable adaptive topology refresh (Lettuce) so failovers are picked up quickly. - 06
Name your connections (
CLIENT SETNAME service:pod) soCLIENT LISTand slowlog entries point to the right service during incidents.
In your stack
- →
Lettuce (Spring Boot's default): Netty-based, thread-safe, one shared connection per node multiplexes all commands; supports sync, async and reactive APIs, Sentinel, Cluster with adaptive topology refresh, and RESP3. Don't wrap it in a large pool unless you use blocking commands or transactions.
- →
Jedis: blocking and simple; a
Jedisinstance isn't thread-safe, so useJedisPoolorJedisPooled/JedisCluster, which pool internally. SizemaxTotalfrom real concurrency and setmaxWaitso pool exhaustion fails fast instead of hanging. - →
Redisson: a higher-level client with distributed locks, semaphores, rate limiters and collections built on Redis. Convenient, but understand what each primitive guarantees (Phase 12) before relying on it.
Code & diagrams
Spring Boot 3.x with Lettuce: explicit timeouts and cluster topology refresh.
spring:
data:
redis:
cluster:
nodes: redis-0:6379,redis-1:6379,redis-2:6379
timeout: 200ms # command timeout
connect-timeout: 1s
client-name: checkout-svc
lettuce:
cluster:
refresh:
adaptive: true # refresh slot map on MOVED/ASK and failures
period: 30simport redis
from redis.backoff import ExponentialWithJitterBackoff
from redis.retry import Retry
r = redis.Redis(
host="redis", port=6379,
socket_connect_timeout=1.0,
socket_timeout=0.2, # 200 ms per command
retry=Retry(ExponentialWithJitterBackoff(base=0.01, cap=0.2), retries=2),
retry_on_error=[redis.ConnectionError, redis.TimeoutError],
health_check_interval=30,
client_name="checkout-svc",
)
# Only idempotent calls should go through automatic retries;
# for INCR/LPUSH use a client with retries=0 and handle errors explicitly.Interview problem
The problem
Redis got slow and took down the checkout service
Redis latency rose from 1 ms to 300 ms for two minutes during a failover. Checkout requests timed out, thread pools filled, health checks failed and Kubernetes restarted pods, making it worse. Redesign the client side so this can't cascade.
You're given
- Java service, Lettuce
- Checkout latency budget 800 ms
- Redis used for cart and rate limiting
The interviewer follows up
Why is a huge connection pool harmful?
When it breaks
No command timeout configured
What you see
A stalled Redis blocks every request thread indefinitely; the app becomes unresponsive and cascades to its callers.
Fix & prevent
Set command and connect timeouts in every client; alert on client-side Redis latency.
Retrying INCR on timeout
What you see
The first attempt actually succeeded; the retry increments again. Counters and quotas drift upward.
Fix & prevent
Never auto-retry non-idempotent commands, or make them idempotent (Lua with a request ID).
Explain it without notes
Compare pooled and multiplexed clients.
Which Redis commands need a dedicated connection in a multiplexed client, and why?
Practice
Configure timeouts and a circuit breaker in your client, then pause Redis with redis-cli DEBUG SLEEP 5 in a lab and observe the app's behaviour.
Trade-offs
- ↔
Short timeouts protect the app but cause more errors during brief blips; pair them with fallbacks.
- ↔
Multiplexing is efficient but concentrates risk on one connection per node.
Done when you can
Every Redis client I configure has timeouts, bounded retries and a named connection.
I can size pools and explain multiplexing trade-offs.
I design fallbacks so a slow Redis can't cascade.