Topic 14.4
Managed Redis and Redis on Kubernetes
In one line
Managed services (ElastiCache, MemoryDB, Memorystore, Azure, Redis Cloud, Upstash) handle failover, patching and backups; running Redis yourself on VMs or Kubernetes gives control and can cost less, but you own everything. Choose based on durability needs, features, engine (Redis or Valkey), cost and team skills.
Think of it like this
Renting a furnished apartment with a building manager versus owning a house. Renting is less work and fixed rules; owning gives control and needs you to fix the roof.
Key ideas
- 01
AWS: ElastiCache (Valkey or Redis OSS engines, cluster mode, Global Datastore, a serverless option) for caching and general use; MemoryDB for durable Redis/Valkey-compatible storage with a multi-AZ transaction log (acknowledged writes survive node loss, at higher write latency). Google Cloud Memorystore (Redis, Redis Cluster and Valkey offerings), Azure Cache for Redis / Azure Managed Redis, Redis Cloud (Redis Ltd., with active-active), and Upstash (serverless, HTTP-accessible, per-request pricing).
- 02
What managed services often restrict: some commands (for example
CONFIG,DEBUG,MODULE), engine versions available, module support, and custom configs. Check that the features you use (Functions, JSON, search,HEXPIRE) exist on your service and engine version. - 03
Kubernetes: Redis is stateful, so use StatefulSets with persistent volumes (or no persistence for pure caches), pod anti-affinity across zones, PodDisruptionBudgets, readiness probes that don't restart pods for transient Redis states, and an operator (such as the OT-Container-Kit redis-operator or vendor operators) or well-maintained Helm charts for Sentinel or Cluster topologies. Cluster mode on Kubernetes needs stable node addresses (
cluster-announce-hostname/IP) because pod IPs change. - 04
Resource settings: memory limit comfortably above
maxmemory(fork and fragmentation headroom); CPU requests realistic; don't CPU-throttle Redis (throttling looks like latency spikes). - 05
Cost thinking: managed pricing is per node-hour with replicas doubling cost; data transfer across AZs adds up for replica traffic and cross-AZ clients. Right-size memory, use reserved capacity for steady workloads, and don't cache what doesn't need caching.
Code & diagrams
# Excerpt: key settings for a self-managed Redis replica set on Kubernetes
apiVersion: apps/v1
kind: StatefulSet
metadata: { name: redis }
spec:
serviceName: redis
replicas: 3
template:
spec:
affinity:
podAntiAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
- labelSelector: { matchLabels: { app: redis } }
topologyKey: topology.kubernetes.io/zone
containers:
- name: redis
image: valkey/valkey:8.1 # pin versions
args: ["--maxmemory", "3gb", "--maxmemory-policy", "noeviction", "--appendonly", "yes"]
resources:
requests: { cpu: "1", memory: "4Gi" }
limits: { memory: "5Gi" } # headroom above maxmemory; no CPU limit
readinessProbe:
exec: { command: ["sh", "-c", "valkey-cli ping | grep PONG"] }
volumeClaimTemplates:
- metadata: { name: data }
spec: { accessModes: ["ReadWriteOnce"], resources: { requests: { storage: 20Gi } } }
---
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata: { name: redis }
spec: { maxUnavailable: 1, selector: { matchLabels: { app: redis } } }Interview problem
The problem
Managed or self-hosted?
A startup runs everything on AWS with a small platform team. They need a cache, a session store and a durable job queue that must not lose acknowledged jobs. Recommend managed vs self-hosted and specific services, with reasons.
The interviewer follows up
Why set a Kubernetes memory limit above maxmemory?
When it breaks
Redis pod with a CPU limit
What you see
CFS throttling pauses the main thread in bursts; latency spikes appear with low average CPU.
Fix & prevent
Set CPU requests, avoid CPU limits for Redis (or give generous ones), and watch throttling metrics.
Explain it without notes
What's the difference between ElastiCache and MemoryDB?
Practice
Compare the monthly cost of a 3-shard, 1-replica cluster with 13 GB nodes on your cloud's managed service vs self-hosted VMs of similar size, including cross-AZ transfer.
Trade-offs
- ↔
Managed: less operational work, fewer knobs and features. Self-hosted: full control, lower infrastructure cost, higher people cost.
Done when you can
I can choose between managed services and self-hosting for a workload.
I know the key Kubernetes settings for running Redis safely.