Architecting High-Performance Distributed Rate Limiting with Redis and Lua in Go
Architecting High-Performance Distributed Rate Limiting with Redis and Lua in Go
Introduction
Modern microservice architectures require strict protection against API abuse, denial-of-service (DoS) attacks, and resource starvation caused by "noisy neighbors." While in-memory rate limiting works well for single-instance monoliths, it fails in distributed environments where client requests are routed across auto-scaling application clusters.
To enforce global API rate limits, systems must rely on a centralized, shared data store. Redis is the industry standard for this pattern due to its low-latency, in-memory data structures. However, a naive implementation involving separate Redis read and write operations introduces race conditions (the "time-of-check to time-of-use" vulnerability) under high concurrency. This article demonstrates how to architect a production-grade, atomic, distributed rate limiter using Go, Redis, and embedded Lua scripting.
Core Concepts & System Architecture
To achieve accurate rate limiting without sacrificing performance, we use the Sliding Window Counter algorithm. Unlike the Fixed Window algorithm—which suffers from burst spikes at boundary resets—the Sliding Window algorithm tracks timestamped requests within a moving time frame.
To make this operation atomic and thread-safe without expensive distributed locks, we leverage Redis Lua scripting. Redis executes Lua scripts in a single-threaded execution context, guaranteeing that no other client can interleave operations during execution. This eliminates race conditions entirely.
Client Request │ ▼ ┌──────────────┐ Lua Script Executed Atomically │ Go API Host ├─────────────────────────────────────────┐ └──────┬───────┘ │ │ (Executes Lua script via EvalSha) ▼ ▼ ┌──────────────┐ ┌──────────────┐ │ Redis Server │ │ Redis Cache │ ◄───────────────────────────────┤ (Sorted Set) │ └──────────────┘ Returns 1 (Allow) or 0 (Block) └──────────────┘
-
Client Layer: Requests enter the system with an identifying key (such as an IP address, API token, or User ID).
-
Middleware Layer: The Go microservice intercepts the request, generates a Redis key based on the identifier, and calls the Lua script.
-
Storage Layer (Redis): Redis stores the timestamps of client requests in a Sorted Set (
ZSET). The Lua script cleans expired logs, counts current requests, and conditionally adds the new timestamp in a single atomic transaction.
Hands-on Implementation
Below is the production-grade implementation of the Sliding Window rate limiter in Go, utilizing the go-redis/v9 client library.
1. The Lua Sliding Window Script
This script runs entirely inside Redis, pruning requests older than the sliding window and validating whether the current request count exceeds the limit.
lua local key = KEYS[1] local now = tonumber(ARGV[1]) local window = tonumber(ARGV[2]) local limit = tonumber(ARGV[3]) local clear_before = now - window
-- Remove elements older than the sliding window redis.call('ZREMRANGEBYSCORE', key, 0, clear_before)
-- Count current requests in the window local current_requests = redis.call('ZCARD', key)
if current_requests < limit then -- Add the current unique timestamp as member and score redis.call('ZADD', key, now, now) -- Set TTL on the key to automatically clean idle clients redis.call('EXPIRE', key, window) return 1 -- Allowed else return 0 -- Rate limited end
2. The Go Rate Limiter Middleware
package main
import (
"context"
"context/metadata"
"errors"
"fmt"
"net/http"
"time"
"github.com/redis/go-redis/v9"
)
const luaScript = local key = KEYS[1]
local now = tonumber(ARGV[1])
local window = tonumber(ARGV[2])
local limit = tonumber(ARGV[3])
local clear_before = now - window
redis.call('ZREMRANGEBYSCORE', key, 0, clear_before)
local current_requests = redis.call('ZCARD', key)
if current_requests < limit then
redis.call('ZADD', key, now, now)
redis.call('EXPIRE', key, window)
return 1
else
return 0
end
type RateLimiter struct { rdb *redis.Client scriptSHA string limit int windowSecs int }
func NewRateLimiter(rdb redis.Client, limit int, windowSecs int) (RateLimiter, error) { ctx := context.Background() sha, err := rdb.ScriptLoad(ctx, luaScript).Result() if err != nil { return nil, fmt.Errorf("failed to load Lua script: %w", err) } return &RateLimiter{ rdb: rdb, scriptSHA: sha, limit: limit, windowSecs: windowSecs, }, nil }
func (rl *RateLimiter) Allow(ctx context.Context, clientID string) (bool, error) { now := time.Now().UnixNano() / int64(time.Millisecond) windowMs := rl.windowSecs * 1000 key := fmt.Sprintf("ratelimit:%s", clientID)
// Try executing with SHA to leverage cached script on Redis
res, err := rl.rdb.EvalSha(ctx, rl.scriptSHA, []string{key}, now, windowMs, rl.limit).Result()
if err != nil {
return false, err
}
allowed, ok := res.(int64)
if !ok {
return false, errors.New("unexpected Redis script output type")
}
return allowed == 1, nil
}
func (rl RateLimiter) Middleware(next http.Handler) http.Handler { return http.HandlerFunc(func(w http.ResponseWriter, r http.Request) { clientID := r.RemoteAddr // Simplify identifying key extraction
allowed, err := rl.Allow(r.Context(), clientID)
if err != nil {
// Fail-open strategy to prevent Redis outages from blocking legitimate users
next.ServeHTTP(w, r)
return
}
if !allowed {
w.Header().Set("Retry-After", fmt.Sprintf("%d", rl.windowSecs))
http.Error(w, "Too Many Requests", http.StatusTooManyRequests)
return
}
next.ServeHTTP(w, r)
})
}
Enterprise Security & Production Hardening
Executing distributed rate limiting in mission-critical architectures requires careful consideration of resiliency and optimization.
Architectural Warning: Rate limiters represent a single point of failure (SPOF) if they fail-closed. In production, always implement a circuit breaker (e.g., using Go's
hystrix-goor standard fallback mechanisms) that switches to a "fail-open" posture if Redis latency exceeds 20ms.
-
Preventing Memory Overrun: The sorted set (
ZSET) can grow excessively if a single attacker floods the API. Always enforce a hard maximum length on the sliding window using a secondary counter, or ensure that Redis is configured with anallkeys-lrueviction policy to discard older keys under low-memory pressure. -
Optimizing Latency via EvalSha: Always load the Lua script once during application bootstrap (
ScriptLoad) and execute requests viaEvalSha. Sending raw scripts on every request wastes network bandwidth and introduces CPU overhead on the Redis server. -
Redis Cluster Considerations: If running in a clustered Redis environment, you must ensure that all keys mapped to a specific client land on the same cluster slot. Accomplish this by using Redis hash tags (e.g.,
{ratelimit:client-123}) to force the hashing algorithm to target the same shard.
Conclusion
By executing sliding window logic atomically inside Redis via Lua scripts, you eliminate race conditions and avoid the heavy performance overhead associated with distributed lock managers. When combined with Go's lightweight concurrency model, this architectural pattern allows systems to process tens of thousands of requests per second with sub-millisecond overhead, protecting core application backends from degradation.
