Command Palette

Search for a command to run...

Hectal
PHASE 7Intermediate ~11 min· topic 3 of 4

Topic 7.3

Designing a Distributed Rate Limiter

In one line

A production limiter combines several scopes (per IP, user, API, tenant), several tiers (free vs premium), atomic Lua decisions, cluster-friendly keys, hot-key handling, and a deliberate choice of what happens when Redis is slow or down.

0/4 · 0%

Think of it like this

Airport security. Every passenger is checked (per user), groups are counted per airline (per tenant), some lanes are for premium travellers with higher throughput (tiers), and if the scanner breaks, the airport has a documented rule: stop everyone, or wave people through with manual checks.

Key ideas

  1. 01

    Where it runs: at the API gateway (Envoy, Kong, NGINX, Spring Cloud Gateway) or as middleware in services. Gateways like Envoy use an external rate-limit service backed by Redis.

  2. 02

    Multiple scopes: each request may be checked against several limits: per IP (abuse), per user (fairness), per API route (protect expensive endpoints), per tenant (plan limits). Check them in one Lua call when keys share a slot, or in a pipeline across slots; reject if any limit is exceeded.

  3. 03

    Tiers: look up the user's plan (cached locally) to pick limits: free 100/min, premium 1,000/min. Store the rules as configuration, not code.

  4. 04

    Scale: 100K requests/sec is fine for a Redis cluster (one Lua call per request, spread over many user keys). Hot keys appear for per-tenant or per-API limits on huge tenants: shard the counter (N sub-keys, each with limit/N) or pre-allocate token batches to each gateway instance (local bucket refilled from Redis).

  5. 05

    Latency budget: the limiter adds one Redis round trip (~0.5 ms). Set tight timeouts (a few ms); a slow limiter must not slow the whole API.

  6. 06

    Failure policy: fail open (allow when Redis is unavailable) keeps the API up but removes protection; fail closed protects the backend but turns a Redis outage into an API outage. Common practice: fail open with a local in-memory fallback limiter (limit ÷ instance count) plus alerting; fail closed only for security-critical limits like login attempts or OTP sends.

  7. 07

    Accuracy vs availability: the limiter is a best-effort control. A few percent overshoot during failover is acceptable; design so it can't take the service down.

Code & diagrams

limiter-architecture.mermaiddiagram
Rendering diagram…
multi-limit.lualua

Check several fixed-window limits in one atomic call; nothing is counted if any limit rejects.

-- KEYS[i] = counter keys (same hash tag); ARGV[i] = limits; ARGV[#KEYS+1] = window sec
local n = #KEYS
local window = tonumber(ARGV[n + 1])
for i = 1, n do
  local c = tonumber(redis.call('GET', KEYS[i]) or '0')
  if c >= tonumber(ARGV[i]) then
    return {0, i}                         -- rejected by limit i
  end
end
for i = 1, n do
  if redis.call('INCR', KEYS[i]) == 1 then
    redis.call('EXPIRE', KEYS[i], window)
  end
end
return {1, 0}
RateLimitFilter.javajava
@Component
public class RateLimitFilter extends OncePerRequestFilter {
    private final RedisRateLimiter redisLimiter;          // wraps the Lua script
    private final LocalRateLimiter fallback;              // e.g. Bucket4j in memory
    private final MeterRegistry metrics;

    @Override
    protected void doFilterInternal(HttpServletRequest req, HttpServletResponse res, FilterChain chain)
            throws IOException, ServletException {
        String user = currentUserId(req);
        Plan plan = plans.forUser(user);                  // cached locally
        Decision d;
        try {
            d = redisLimiter.check(user, req.getRequestURI(), plan);   // ~0.5 ms, 5 ms timeout
        } catch (RedisConnectionFailureException | QueryTimeoutException e) {
            metrics.counter("ratelimit.redis_error").increment();
            d = fallback.check(user, plan);               // fail open, but still bounded locally
        }
        if (!d.allowed()) {
            res.setStatus(429);
            res.setHeader("Retry-After", String.valueOf(d.retryAfterSeconds()));
            return;
        }
        res.setHeader("X-RateLimit-Remaining", String.valueOf(d.remaining()));
        chain.doFilter(req, res);
    }
}

Interview problem

The problem

API rate limiter: tiers, scopes, 100K rps, node failure

Free users get 100 requests/minute and premium users 1,000. Then add limits per IP, per user, per API and per tenant. Then scale to 100K requests/sec. Then a Redis node fails. Walk through the design.

You're given

  • Free 100/min, premium 1,000/min
  • Per IP, user, API, tenant
  • 100K rps
  • Redis Cluster with replicas

The interviewer follows up

01

Why not keep all counters in each gateway's memory?

02

Would you use Redis replicas to read limiter state?

When it breaks

Limiter fails closed on a Redis timeout

What you see

A 30-second Redis failover becomes a 30-second outage of every API, even though the backends are healthy.

Fix & prevent

Fail open with a local fallback for normal limits, and alert; fail closed only where abuse risk outweighs availability.

Limiter call without a timeout

What you see

When Redis is slow, every request waits for it; API latency and thread pools blow up.

Fix & prevent

Aggressive client timeouts (few ms) and a circuit breaker around the limiter.

Explain it without notes

01

Explain fail-open vs fail-closed for a rate limiter and when you'd choose each.

Practice

01

Add per-IP and per-user limits to one endpoint, then stop Redis and confirm the fallback works and emits a metric.

Trade-offs

  • ↔

    Global accuracy (Redis) vs availability and latency (local); most designs combine both.

  • ↔

    More scopes give better protection but add Redis calls per request.

Done when you can

  • I can design a multi-scope, multi-tier limiter with atomic decisions.

  • I can handle hot keys, node failures and fail-open/fail-closed choices.