Back to articles
Technology Insight

Architecting High-Performance Distributed Rate Limiting with Redis and Lua in Go

August 14, 2026

Architecting High-Performance Distributed Rate Limiting with Redis and Lua in Go

Introduction

Modern microservice architectures require strict protection against API abuse, denial-of-service (DoS) attacks, and resource starvation caused by "noisy neighbors." While in-memory rate limiting works well for single-instance monoliths, it fails in distributed environments where client requests are routed across auto-scaling application clusters.

To enforce global API rate limits, systems must rely on a centralized, shared data store. Redis is the industry standard for this pattern due to its low-latency, in-memory data structures. However, a naive implementation involving separate Redis read and write operations introduces race conditions (the "time-of-check to time-of-use" vulnerability) under high concurrency. This article demonstrates how to architect a production-grade, atomic, distributed rate limiter using Go, Redis, and embedded Lua scripting.

Core Concepts & System Architecture

To achieve accurate rate limiting without sacrificing performance, we use the Sliding Window Counter algorithm. Unlike the Fixed Window algorithm—which suffers from burst spikes at boundary resets—the Sliding Window algorithm tracks timestamped requests within a moving time frame.

To make this operation atomic and thread-safe without expensive distributed locks, we leverage Redis Lua scripting. Redis executes Lua scripts in a single-threaded execution context, guaranteeing that no other client can interleave operations during execution. This eliminates race conditions entirely.

Client Request │ ▼ ┌──────────────┐ Lua Script Executed Atomically │ Go API Host ├─────────────────────────────────────────┐ └──────┬───────┘ │ │ (Executes Lua script via EvalSha) ▼ ▼ ┌──────────────┐ ┌──────────────┐ │ Redis Server │ │ Redis Cache │ ◄───────────────────────────────┤ (Sorted Set) │ └──────────────┘ Returns 1 (Allow) or 0 (Block) └──────────────┘

  1. Client Layer: Requests enter the system with an identifying key (such as an IP address, API token, or User ID).

  2. Middleware Layer: The Go microservice intercepts the request, generates a Redis key based on the identifier, and calls the Lua script.

  3. Storage Layer (Redis): Redis stores the timestamps of client requests in a Sorted Set (ZSET). The Lua script cleans expired logs, counts current requests, and conditionally adds the new timestamp in a single atomic transaction.

Hands-on Implementation

Below is the production-grade implementation of the Sliding Window rate limiter in Go, utilizing the go-redis/v9 client library.

1. The Lua Sliding Window Script

This script runs entirely inside Redis, pruning requests older than the sliding window and validating whether the current request count exceeds the limit.

lua local key = KEYS[1] local now = tonumber(ARGV[1]) local window = tonumber(ARGV[2]) local limit = tonumber(ARGV[3]) local clear_before = now - window

-- Remove elements older than the sliding window redis.call('ZREMRANGEBYSCORE', key, 0, clear_before)

-- Count current requests in the window local current_requests = redis.call('ZCARD', key)

if current_requests < limit then -- Add the current unique timestamp as member and score redis.call('ZADD', key, now, now) -- Set TTL on the key to automatically clean idle clients redis.call('EXPIRE', key, window) return 1 -- Allowed else return 0 -- Rate limited end

2. The Go Rate Limiter Middleware

package main
import (
    "context"
    "context/metadata"
    "errors"
    "fmt"
    "net/http"
    "time"
"github.com/redis/go-redis/v9"

)

const luaScript = local key = KEYS[1] local now = tonumber(ARGV[1]) local window = tonumber(ARGV[2]) local limit = tonumber(ARGV[3]) local clear_before = now - window redis.call('ZREMRANGEBYSCORE', key, 0, clear_before) local current_requests = redis.call('ZCARD', key) if current_requests < limit then redis.call('ZADD', key, now, now) redis.call('EXPIRE', key, window) return 1 else return 0 end

type RateLimiter struct { rdb *redis.Client scriptSHA string limit int windowSecs int }

func NewRateLimiter(rdb redis.Client, limit int, windowSecs int) (RateLimiter, error) { ctx := context.Background() sha, err := rdb.ScriptLoad(ctx, luaScript).Result() if err != nil { return nil, fmt.Errorf("failed to load Lua script: %w", err) } return &RateLimiter{ rdb: rdb, scriptSHA: sha, limit: limit, windowSecs: windowSecs, }, nil }

func (rl *RateLimiter) Allow(ctx context.Context, clientID string) (bool, error) { now := time.Now().UnixNano() / int64(time.Millisecond) windowMs := rl.windowSecs * 1000 key := fmt.Sprintf("ratelimit:%s", clientID)

// Try executing with SHA to leverage cached script on Redis
res, err := rl.rdb.EvalSha(ctx, rl.scriptSHA, []string{key}, now, windowMs, rl.limit).Result()
if err != nil {
    return false, err
}

allowed, ok := res.(int64)
if !ok {
    return false, errors.New("unexpected Redis script output type")
}

return allowed == 1, nil

}

func (rl RateLimiter) Middleware(next http.Handler) http.Handler { return http.HandlerFunc(func(w http.ResponseWriter, r http.Request) { clientID := r.RemoteAddr // Simplify identifying key extraction

    allowed, err := rl.Allow(r.Context(), clientID)
    if err != nil {
        // Fail-open strategy to prevent Redis outages from blocking legitimate users
        next.ServeHTTP(w, r)
        return
    }

    if !allowed {
        w.Header().Set("Retry-After", fmt.Sprintf("%d", rl.windowSecs))
        http.Error(w, "Too Many Requests", http.StatusTooManyRequests)
        return
    }

    next.ServeHTTP(w, r)
})

}

Enterprise Security & Production Hardening

Executing distributed rate limiting in mission-critical architectures requires careful consideration of resiliency and optimization.

Architectural Warning: Rate limiters represent a single point of failure (SPOF) if they fail-closed. In production, always implement a circuit breaker (e.g., using Go's hystrix-go or standard fallback mechanisms) that switches to a "fail-open" posture if Redis latency exceeds 20ms.

  • Preventing Memory Overrun: The sorted set (ZSET) can grow excessively if a single attacker floods the API. Always enforce a hard maximum length on the sliding window using a secondary counter, or ensure that Redis is configured with an allkeys-lru eviction policy to discard older keys under low-memory pressure.

  • Optimizing Latency via EvalSha: Always load the Lua script once during application bootstrap (ScriptLoad) and execute requests via EvalSha. Sending raw scripts on every request wastes network bandwidth and introduces CPU overhead on the Redis server.

  • Redis Cluster Considerations: If running in a clustered Redis environment, you must ensure that all keys mapped to a specific client land on the same cluster slot. Accomplish this by using Redis hash tags (e.g., {ratelimit:client-123}) to force the hashing algorithm to target the same shard.

Conclusion

By executing sliding window logic atomically inside Redis via Lua scripts, you eliminate race conditions and avoid the heavy performance overhead associated with distributed lock managers. When combined with Go's lightweight concurrency model, this architectural pattern allows systems to process tens of thousands of requests per second with sub-millisecond overhead, protecting core application backends from degradation.