Building a Distributed API Gateway with Caddy Server and Redis Dynamic Rate Limiting across Multi-VPS Infrastructure
Introduction to Modern API Management
In contemporary enterprise architecture, modernizing the microservices ecosystem requires a robust, scalable, and highly available entry point. As applications grow from single-server deployments to multi-VPS (Virtual Private Server) environments, managing cross-cutting concerns like traffic routing, SSL/TLS termination, and security policy enforcement becomes increasingly complex. This is where an API Gateway becomes indispensable.
While legacy solutions like Nginx or HAProxy have traditionally dominated this space, Caddy Server has emerged as a formidable alternative. Known for its automatic HTTPS management, modular architecture, and developer-friendly configuration syntax, Caddy can easily be extended into a distributed API Gateway. Combined with Redis for centralized state management, Caddy provides an elegant, high-performance solution for dynamic rate limiting across clustered infrastructure.
The Architecture: Multi-VPS and Distributed Gateway
Operating a production-grade infrastructure across multiple VPS instances mitigates the risk of a single point of failure (SPOF) and improves regional latency. However, a distributed gateway layer introduces a significant challenge: state synchronization. If Rate Limiting is managed locally on each individual VPS, a malicious actor could bypass thresholds by simply distributing requests across different gateway nodes.
To solve this, our architecture utilizes a centralized Redis cluster or a highly available Redis sentinel setup. Each Caddy instance acts as a stateless gateway node, inspecting incoming traffic and querying the shared Redis backend in real-time to track and enforce request quotas uniformly across the entire cluster.
Key Architectural Components:
- Load Balancer (Layer 4): Distributes incoming TCP/HTTPS traffic evenly among available Caddy VPS nodes.
- Caddy Edge Nodes (Layer 7 API Gateways): Handle TLS termination, reverse proxy routing, header manipulation, and initial request filtering.
- Redis Cache Cluster: Acts as the centralized, high-speed memory store tracking client request counters and time-to-live (TTL) states.
- Upstream Services: The backend microservices running across your infrastructure that need protection and routing.
Step-by-Step Guide: Configuring Caddy with Redis Rate Limiting
To implement dynamic rate limiting based on a shared Redis instance, we need to extend standard Caddy functionality using custom modules. The standard distribution of Caddy does not include a Redis rate-limiting plugin natively, so we utilize xcaddy, Caddy's official build tool, to compile a custom binary containing the necessary modules.
Step 1: Compiling Caddy with the Redis Rate Limit Module
First, ensure Go is installed on your development or build machine. Execute the following command to build a Caddy binary equipped with the layer4 and rate-limiting capabilities:
xcaddy build
--with [github.com/mholt/caddy-ratelimit](https://github.com/mholt/caddy-ratelimit)Note: In production environments, always ensure you pin your module versions to specific semantic tags to guarantee reproducible builds across your multi-VPS automated deployment pipelines.
Step 2: Designing the Global Caddyfile Configuration
Once you have deployed the custom Caddy binary to all your VPS nodes, create a centralized Caddyfile configuration. Below is a production-ready configuration blueprint tailored for a distributed microservices gateway:
{
# Global configuration options
admin 0.0.0.0:2019
on_demand_tls {
ask http://localhost:8080/allowed_domains
}
}
api.yourcompany.com {
# Enable compression for modern clients
encode gzip zstd
# Security Headers
header {
X-XSS-Protection "1; mode=block"
X-Content-Type-Options "nosniff"
X-Frame-Options "DENY"
Referrer-Policy "strict-origin-when-cross-origin"
Strict-Transport-Security "max-age=63072000; includeSubDomains; preload"
}
# Route definitions with dynamic rate limiting
handle /v1/auth/* {
# Stricter limits for authentication endpoints
rate_limit {
zone auth_limit {
key {remote_ip}
window 1m
max_requests 10
}
redis {
address "redis-cluster.internal:6379"
password "your_secure_redis_password"
db 0
}
}
reverse_proxy [http://auth-service.internal:8081](http://auth-service.internal:8081)
}
handle /v1/public/* {
# Generous limits for public content
rate_limit {
zone public_limit {
key {remote_ip}
window 1s
max_requests 100
}
redis {
address "redis-cluster.internal:6379"
password "your_secure_redis_password"
db 0
}
}
reverse_proxy [http://public-service.internal:8082](http://public-service.internal:8082)
}
# Fallback for unmatched routes
handle {
respond "Not Found" 404
}
}Implementing Dynamic Rate Limiting Rules
Static rate limits defined in configuration files are useful but often insufficient for modern SaaS architectures. For advanced implementations, you may want to enforce dynamic rate limiting based on specific client attributes, such as API Keys, User Roles, or Subscription Tiers (e.g., Free vs. Premium).
By leveraging Caddy's flexible expressions and placeholders, you can extract JWT payload claims or custom headers (like Authorization or X-API-Key) and pass them as the rate-limiting identifier key. Instead of mapping solely to {remote_ip}, modifying your configuration key to {http.request.header.X-API-Key} allows your distributed gateways to enforce quotas globally, regardless of whether the user switches networks or utilizes proxies.
Example: Tier-Based Routing Configuration
Consider a scenario where you want to prioritize traffic and apply divergent rate limits based on user tiers. You can split traffic elegantly using Caddy matchers:
@premium_users {
header X-User-Tier Premium
}
@free_users {
header X-User-Tier Free
}
# Handle Free Tier traffic
handle @free_users {
rate_limit {
zone free_zone {
key {http.request.header.X-API-Key}
window 1h
max_requests 1000
}
redis { address "redis-cluster.internal:6379" }
}
reverse_proxy [http://backend-cluster.internal](http://backend-cluster.internal)
}
# Handle Premium Tier traffic
handle @premium_users {
rate_limit {
zone premium_zone {
key {http.request.header.X-API-Key}
window 1h
max_requests 50000
}
redis { address "redis-cluster.internal:6379" }
}
reverse_proxy [http://backend-cluster.internal](http://backend-cluster.internal)
}Operational Best Practices for Multi-VPS Gateways
Deploying and operating a distributed API Gateway structure seamlessly requires rigorous operational focus. Below are essential strategies to ensure long-term stability and resilience:
- Implement Redis Connection Pooling: Ensure your Caddy configuration reuse Redis TCP connections efficiently to prevent socket exhaustion under heavy concurrent request spikes.
- Graceful Degradation (Fail-Open vs. Fail-Closed): Decide how your API Gateway behaves if the centralized Redis cluster goes offline. In high-availability business scenarios, a fail-open strategy (allowing traffic through with local fallbacks) is typically preferred over completely denying service to users.
- Centralized Log Collection: Enable structured JSON logging within Caddy. Stream these logs to an external aggregate stack (such as ELK, Grafana Loki, or Datadog) to gain immediate, comprehensive observability across all active VPS locations.
- Automated CI/CD Deployments: Treat your Caddyfiles as infrastructure code. Store configurations in Git repositories, run automated validation checks via
caddy validate, and roll out changes progressively using zero-downtime reloads (caddy reload).
Conclusion
Transforming Caddy Server into a distributed API gateway paired with dynamic, Redis-backed rate limiting offers a formidable combination of speed, simplicity, and elite scalability for multi-VPS environments. By moving away from monolithic gateway frameworks, enterprise engineering teams can harness Caddy's automatic certificate lifecycle workflows and memory efficiency while ensuring precise, real-time security enforcement across all entry points. Implement this architecture today to elevate the resilience and security profile of your backend ecosystem.
