Command Palette

Search for a command to run...

Hectal
PHASE 13Advanced ~9 min· topic 4 of 5

Topic 13.4

Monitoring Redis: INFO, SLOWLOG, LATENCY, and Alerts

In one line

Monitor Redis with a small set of metrics that map to real failures: memory and evictions, hit ratio, ops and latency, clients and blocked clients, persistence status, replication links and lag, and cluster state. Collect them with an exporter, graph them, and alert on symptoms users feel.

0/5 · 0%

Think of it like this

A car dashboard. You don't need every sensor reading, just speed, fuel, temperature and warning lights. The skill is knowing which light means "pull over now" and which means "service soon".

Key ideas

  1. 01

    Sources: INFO sections (server, clients, memory, persistence, stats, replication, cpu, commandstats, keyspace, cluster), SLOWLOG (commands slower than slowlog-log-slower-than, default 10,000 µs, keeping slowlog-max-len entries), LATENCY (events over latency-monitor-threshold ms: fork, AOF fsync, expire cycles, commands), MEMORY STATS/MEMORY DOCTOR, and CLIENT LIST.

  2. 02

    Key metrics: used_memory vs maxmemory and RSS; mem_fragmentation_ratio; evicted_keys and expired_keys rates; keyspace_hits/keyspace_misses (hit ratio); instantaneous_ops_per_sec and per-command rates; connected_clients, blocked_clients, rejected_connections; rdb_last_bgsave_status, aof_last_write_status, latest_fork_usec; master_link_status, replica lag and offsets; cluster_state, cluster_slots_fail.

  3. 03

    Client-side metrics matter as much: command latency percentiles as your app sees them, timeouts, connection pool wait time and errors. Server-side slowlog misses network and queueing time.

  4. 04

    Tooling: the Prometheus redis_exporter (oliver006) with Grafana dashboards, managed-service metrics in CloudWatch/Azure Monitor/Cloud Monitoring, and RedisInsight for ad-hoc analysis.

  5. 05

    Alerts that matter: p99 latency (client side) above SLO; memory above ~85% of maxmemory; evictions on a non-cache instance; hit ratio drop; replication link down or lag above threshold; persistence failures; cluster_state not ok; rejected connections; big increases in slowlog entries.

Code & diagrams

diagnostics.redisredis
127.0.0.1:6379> CONFIG SET latency-monitor-threshold 100
OK
127.0.0.1:6379> LATENCY LATEST
1) 1) "fork"
   2) (integer) 1727520000
   3) (integer) 412              # latest ms
   4) (integer) 890              # max ms
2) 1) "command"
   2) (integer) 1727519800
   3) (integer) 230
   4) (integer) 2310
127.0.0.1:6379> LATENCY DOCTOR
Dave, I have observed latency spikes in this Redis instance...
1. fork: 3 latency spikes (average 511ms, worst 890ms)
- Consider disabling transparent huge pages...
127.0.0.1:6379> SLOWLOG GET 2
1) 1) (integer) 1204  2) (integer) 1727519800  3) (integer) 230112
   4) 1) "ZRANGE" 2) "lb:global" 3) "0" 4) "-1"
   5) "10.0.3.7:51022"  6) "leaderboard-api"
127.0.0.1:6379> INFO clients
connected_clients:1204
blocked_clients:18
maxclients:10000
alerts.ymlyaml

Prometheus rules using redis_exporter metric names.

groups:
- name: redis
  rules:
  - alert: RedisMemoryHigh
    expr: redis_memory_used_bytes / redis_memory_max_bytes > 0.85
    for: 10m
  - alert: RedisEvictionsOnStatefulInstance
    expr: rate(redis_evicted_keys_total{role="state"}[5m]) > 0
  - alert: RedisReplicationBroken
    expr: redis_connected_slaves < 1
    for: 2m
  - alert: RedisHitRatioDrop
    expr: |
      rate(redis_keyspace_hits_total[5m])
        / (rate(redis_keyspace_hits_total[5m]) + rate(redis_keyspace_misses_total[5m])) < 0.8
    for: 15m
  - alert: RedisRejectedConnections
    expr: increase(redis_rejected_connections_total[5m]) > 0

Interview problem

The problem

Build the Redis dashboard for an on-call team

Design the dashboard and alerts for a Redis Cluster serving sessions and caches. The on-call engineer should be able to tell in one minute whether Redis is the cause of an incident.

The interviewer follows up

01

Why isn't a green Redis dashboard enough to clear Redis in an incident?

Explain it without notes

01

Name eight Redis metrics you'd alert on and why.

Practice

01

Run redis_exporter against your lab and build a Grafana panel for hit ratio and evictions.

Trade-offs

  • ↔

    Lower slowlog and latency thresholds catch more problems but add noise; start at 10 ms and 100 ms and adjust.

Done when you can

  • I know which INFO fields matter and what SLOWLOG and LATENCY show.

  • I can build dashboards and alerts that tie Redis health to user impact.