Back to articles
Technology Insight

Building a Distributed API Gateway: Dynamic Rate Limiting with Caddy Server and Redis

May 30, 2026

Introduction to Modern API Management

In the era of microservices and cloud-native architectures, managing traffic efficiently is paramount. An API Gateway serves as the single entry point for all clients, handling critical cross-cutting concerns such as routing, SSL termination, authentication, and traffic shaping. While traditional gateways like Nginx, Kong, or Envoy dominate the landscape, Caddy Server has emerged as a formidable challenger.

Caddy is a modern, extensible, open-source web server written in Go. Famous for its automatic HTTPS management, Caddy also features a powerful, modular architecture that makes it an excellent choice for a distributed API Gateway. However, in a distributed topology where multiple gateway instances run behind a global load balancer, isolated local traffic management is insufficient. To enforce consistent, real-time rate limiting across an entire cluster, a shared memory layer is required. This guide explores how to configure Caddy Server as a distributed API Gateway featuring dynamic rate limiting backed by Redis.

The Architecture: Distributed Gateway with Shared Memory

When scaling out API gateways horizontally, relying on local, in-memory rate limiting introduces a fundamental flaw: traffic spikes can bypass limits if a client's requests are distributed evenly across different gateway nodes. For instance, if a user is limited to 60 requests per minute and you have three gateway instances, local limiting might allow that user to make up to 180 requests per minute.

To solve this, we introduce a centralized state management layer using Redis. Redis provides an ultra-low latency, in-memory data structure store that perfectly matches the high-throughput requirements of an API gateway. By utilizing the Token Bucket or Leaky Bucket algorithms implemented via Redis scripts or modules, all Caddy instances can check and decrement rate limits globally and atomically.

Key Benefits of this Architecture

  • Global State Consistency: Rate limits are enforced accurately across all geographic regions and gateway instances.
  • Dynamic Configuration: Limits can be adjusted in real-time within Redis without restarting the Caddy instances.
  • High Performance: Redis operates in-memory, ensuring that the rate-limiting check adds sub-millisecond latency to the request lifecycle.
  • Memory Efficiency: Sliding window logs or token counters use minimal data footprints inside Redis.
---

Step 1: Customizing Caddy with XCaddy

Out of the box, Caddy includes standard reverse proxy capabilities but requires external modules for advanced rate limiting and Redis integration. Because Caddy is compiled in Go, adding features requires building a custom binary. Fortunately, the Caddy team provides a tool called xcaddy to automate this process seamlessly.

To build a Caddy binary that includes the necessary HTTP rate-limiting modules with Redis support, execute the following command in your terminal:

xcaddy build --with [github.com/mholt/caddy-ratelimit](https://github.com/mholt/caddy-ratelimit)

Note: Ensure you verify the specific third-party module repository that aligns with your exact Redis architecture requirements, such as support for Redis Cluster or Redis Sentinel.

---

Step 2: Designing the Caddyfile Configuration

Caddy's configuration file, known as the Caddyfile, is celebrated for its human-readable and expressive syntax. Below is an enterprise-grade configuration mapping a distributed API gateway setup with integrated Redis-backed rate limiting.

{
    order ratelimit before reverse_proxy
}

api.yourcompany.com {
    # Global logging for telemetry and analytics
    log {
        output file /var/log/caddy/api_gateway.log
        format json
    }

    # Dynamic Rate Limiting Layer
    ratelimit {
        zone global_api {
            key {http.request.remote}
            window 1m
            max_events 100
            
            # Redis Backend Integration
            storage redis {
                address "redis-cluster.internal:6379"
                password "your_secure_redis_password"
                db 0
                timeout 500ms
            }
        }
    }

    # Microservices Routing Logic
    handle /v1/users* {
        reverse_proxy [http://user-service.internal:8081](http://user-service.internal:8081)
    }

    handle /v1/orders* {
        reverse_proxy [http://order-service.internal:8082](http://order-service.internal:8082)
    }

    # Fallback for unmapped routes
    handle {
        respond "Route not found" 404
    }
}

Anatomy of the Configuration

Let's dissect the critical components of this configuration to understand how it operates under heavy production loads:

  1. The Order Directive: Caddy evaluates middleware chains strictly. By utilizing order ratelimit before reverse_proxy, we guarantee that unauthorized or excessive traffic is rejected immediately at the edge, saving internal backend computing resources.
  2. The Rate Limit Zone: We define a zone named global_api. The key directive specifies how clients are identified. While {http.request.remote} uses the client IP, this can be dynamically changed to an HTTP header value like {http.request.header.Authorization} for authenticated API consumers.
  3. Storage Configuration: This instructs the plugin to bypass local memory and utilize the Redis instance located at redis-cluster.internal:6379. The 500ms timeout ensures that if the Redis cluster encounters a catastrophic failure, the gateway will fail-open or fail-closed gracefully rather than hanging indefinitely.
---

Step 3: Implementing Dynamic Tiered Limits

A static rate limit applied uniformly across all endpoints rarely satisfies modern business requirements. Production API Gateways usually demand tiered rate limiting based on subscription levels (e.g., Free, Silver, Gold). Achieving this dynamically involves combining Caddy's request evaluation expressions with Redis keys.

By reading the client's tier from a JWT token or an early authentication middleware, Caddy can dynamically evaluate variables to determine which rate limit zone or threshold to apply. For example, using Caddy's advanced JSON configuration (or custom Go plugins), you can evaluate the request context and route the client to differentiated limits stored globally in Redis.

Example: Handling Rejections Professionally

When a client exceeds their allocated quota, the gateway must return clear, descriptive error payloads along with the standard HTTP 429 Too Many Requests status code. It is highly recommended to append standard headers such as Retry-After to inform client applications exactly when they can safely resume making requests, reducing unnecessary programmatic retry spam against your architecture.

---

Production Operational Considerations

Deploying Caddy and Redis as your core API infrastructure requires strict adherence to operational best practices to guarantee resilience and high availability:

  • Redis High Availability: Do not rely on a single Redis node. Utilize Redis Sentinel or a native Redis Cluster with replication to prevent the rate limiting storage from becoming a single point of failure (SPOF).
  • Circuit Breaking: Implement a fallback mechanism within Caddy. If the Redis storage cluster goes offline, Caddy should instantly fallback to localized in-memory rate limiting to maintain uptime, generating high-severity alerts to your DevOps engineering teams.
  • Monitoring and Telemetry: Export Caddy metrics using its built-in Prometheus endpoint. Monitor the latency of the rate limiting middleware alongside Redis command execution metrics to spot network bottlenecks early.

Conclusion

Transitioning from traditional API management platforms to a streamlined, distributed Caddy Server + Redis gateway empowers engineering teams with incredible performance, lower infrastructure overhead, and absolute control over traffic shaping. By offloading rate limits to a centralized Redis cluster, you ensure consistent, deterministic protection for your microservices. As your infrastructure demands expand, this cloud-native architecture scales effortlessly alongside your user base.

Building a Distributed API Gateway: Dynamic Rate Limiting with Caddy Server and Redis | DPTCloud