Scaling Caddy Server: How to Efficiently Manage 10,000+ Custom Customer Domains on a Single VPS Cluster
Introduction: The Multi-Tenant Custom Domain Challenge
For modern Software-as-a-Service (SaaS) platforms, e-commerce builders, and website builders, allowing users to point their custom domains (e.g., [www.customer-brand.com](https://www.customer-brand.com)) to a centralized infrastructure is a core requirement. However, managing SSL/TLS certificates and routing for tens of thousands of external domains simultaneously introduces significant engineering hurdles.
Traditionally, infrastructure teams had to rely on complex combinations of Nginx, external Lua scripts, and cron jobs to manage Let's Encrypt certificates. This approach often leads to configuration bloat, reload delays, and frequent rate-limiting issues. Enter Caddy Server—a modern, memory-safe web server written in Go that handles automatic TLS configuration natively. But how does Caddy scale when pushed to the limit of 10,000+ custom domains running on a lean, cost-effective VPS cluster? This comprehensive guide explores the architectural adjustments, configuration tweaks, and kernel optimizations required to achieve seamless multi-tenant domain routing at scale.
1. Leveraging Caddy's Secret Weapon: On-Demand TLS
At the heart of scaling thousands of arbitrary user domains is Caddy’s On-Demand TLS capability. Instead of pre-generating certificates for 10,000 domains at startup (which would crash the server and trigger immediate CA rate limits), Caddy generates certificates on the fly during the initial TLS handshake when a user visits the domain for the first time.
The Risks of Unrestricted On-Demand TLS
If left unconfigured, On-Demand TLS opens your cluster up to serious Distributed Denial of Service (DDoS) and resource exhaustion attacks. A malicious actor could point thousands of wildcard domains to your VPS IP address, forcing Caddy to request certificates for fake domains, rapidly exhausting your Let's Encrypt quotas and filling up your disk space.
Implementing an Ask Endpoint Authentication Pattern
To mitigate this risk, Caddy requires an internal ask endpoint. Before Caddy requests an SSL certificate from an ACME provider (like Let's Encrypt or ZeroSSL), it sends a quick HTTP GET request to your internal backend API to verify if the domain is registered and allowed on your platform.
# Example API Verification logic
GET /check-domain?domain=customer-brand.com
# Response: 200 OK (Allowed) or 404 Not Found (Denied)
2. Production-Grade Caddyfile Configuration
To handle a massive influx of diverse traffic smoothly, you must move away from standard default settings. Below is an optimized Caddyfile architecture specifically tuned for enterprise-grade multi-tenancy and high-volume reverse proxying:
{
# Global options for scaling
on_demand_tls {
ask [http://127.0.0.1:8000/api/v1/domains/validate](http://127.0.0.1:8000/api/v1/domains/validate)
interval 2m
burst 5
}
# Distributed storage configuration for VPS clustering
storage redis {
host "10.0.0.5"
port 6379
db 1
timeout 5s
}
# Performance Tuning
servers {
timeouts {
read_idle 30s
read_header 10s
write 30s
idle 2m
}
}
}
# Dynamic Multi-Tenant Reverse Proxy Handler
:443 {
tls {
on_demand
}
# Security Headers to protect downstream applications
header {
X-XSS-Protection "1; mode=block"
X-Content-Type-Options "nosniff"
X-Frame-Options "DENY"
Referrer-Policy "strict-origin-when-cross-origin"
Strict-Transport-Security "max-age=31536000; includeSubDomains; preload"
}
# Compression optimization
encode gzip zstd
# Route traffic to the internal application backend cluster
reverse_proxy 10.0.0.10:8080 10.0.0.11:8080 {
lb_policy random_choose 2
health_uri /healthz
health_interval 10s
health_timeout 3s
transport http {
dial_timeout 2s
keepalive 30s
tls_insecure_skip_verify
}
}
}
3. Setting Up a Distributed Storage Backend
When running a cluster of multiple VPS nodes behind a Layer 4 Load Balancer, local file storage for TLS certificates is no longer viable. If Node A fetches a certificate for a domain, Node B must have immediate access to it if a subsequent request hits that second server. If they do not share storage, both nodes will issue duplicate certificate requests, triggering severe Let's Encrypt rate limits.
To resolve this, you should deploy a distributed storage plugin like caddy-dns/redis or an S3-compatible backend. This guarantees that:
- ACME Challenges are Unified: All nodes write and read lock states simultaneously.
- Immediate Failover: If one VPS node fails, alternative nodes seamlessly fetch existing certificates from the central Redis/S3 layer without downtime.
- Reduced Disk I/O: High-speed RAM caching through Redis drastically cuts down on local SSD wear-and-tear.
4. Linux Kernel Tuning for High-Concurrency Network I/O
Even the most optimized Caddy configuration will bottleneck if the underlying Linux kernel on your VPS cluster isn't tailored for ultra-high concurrency. By default, Linux is optimized for general desktop or low-volume server usage. To process thousands of simultaneous connections without dropping packets, append the following parameters to your /etc/sysctl.conf file:
System Administrator Note: Always backup your sysctl parameters before applying modifications to a live production environment.
# Increase the maximum number of open files and file descriptors
fs.file-max = 2097152
# Maximize network connection backlog queue
net.core.somaxconn = 65535
net.core.netdev_max_backlog = 50000
# Optimize local port ranges for outbound proxy connections
net.ipv4.ip_local_port_range = 1024 65535
# Enable TCP time-wait reuse for rapid connection recycling
net.ipv4.tcp_tw_reuse = 1
# Adjust maximum TCP buffer memory sizes
net.core.rmem_max = 16777216
net.core.wmem_max = 16777216
net.ipv4.tcp_rmem = 4096 87380 16777216
net.ipv4.tcp_wmem = 4096 65536 16777216
After saving the configurations, run sudo sysctl -p to instantly apply the optimizations without requiring a full system reboot.
Adjusting Security Limits (ulimit)
Additionally, modify /etc/security/limits.conf to prevent Caddy from running out of file descriptors when open socket connections spike during peak traffic hours:
caddy soft nofile 500000
caddy hard nofile 500000
5. Monitoring, Observability, and Proactive Alerting
Operating a multi-tenant custom domain router at scale demands absolute visibility into system health. Caddy provides native integration with Prometheus via its internal metrics endpoint.
Key metrics that your DevOps team must actively monitor include:
caddy_tls_certificates_total: Tracks the total number of managed certificates currently active across the system.caddy_http_request_duration_seconds: Monitors latency spikes indicating backend exhaustion or upstream network bottlenecks.caddy_tls_issuance_failed_total: Directly isolates failure trends in the On-Demand TLS pipeline, highlighting issues with customer DNS setups or rate limits.
Conclusion
Scaling Caddy Server to efficiently handle over 10,000 custom client domains on a standard VPS cluster is entirely achievable when configured correctly. By implementing a highly secure On-Demand TLS Ask Endpoint, migrating to a centralized Redis storage layer, and executing fine-grained Linux kernel tuning, you can transform Caddy into an resilient, automated reverse-proxy powerhouse. This setup not only guarantees exceptional performance for your users but also minimizes infrastructure overhead, proving that enterprise-level infrastructure doesn't require an enterprise-level budget.
