Scaling On-Demand TLS: Optimizing Caddy Server v3 for 50,000+ Customer Domains
Introduction: The Multi-Tenant Custom Domain Challenge
In modern Software-as-a-Service (SaaS) and multi-tenant architectures, allowing users to point their custom domains to your platform is a standard requirement. However, managing SSL/TLS certificates at scale—specifically for 50,000 or more custom domains—presents a massive infrastructure challenge. Traditionally, this required complex automated scripts, cron jobs, and fragile Nginx configurations coupled with Certbot.
Enter Caddy Server v3. Known for its production-ready, default-HTTPS nature, Caddy simplifies this complexity through its native On-Demand TLS feature. Instead of pre-generating certificates for thousands of domains, Caddy generates them on-the-fly during the TLS handshake. While this sounds like a silver bullet, running On-Demand TLS at a scale of 50,000+ domains requires meticulous optimization to prevent rate limiting, high latency, and single points of failure. This comprehensive guide details how to optimize Caddy v3 for enterprise-grade custom domain management.
Understanding Caddy's On-Demand TLS Mechanism
Before optimizing, it is critical to understand the lifecycle of an On-Demand TLS request. When an un-cached domain hits your Caddy instance:
- TCP Connection & Client Hello: The client initiates a TLS handshake specifying the custom domain via Server Name Indication (SNI).
- Ask Endpoint Check: Caddy intercepts the handshake and sends an internal HTTP request to an internal API (the "Ask" endpoint) to verify if the domain is authorized.
- ACME Challenge: If authorized, Caddy checks its local/distributed cache. If the certificate is missing, it automatically triggers an ACME challenge (via Let's Encrypt or ZeroSSL) to issue the certificate.
- Handshake Completion: The certificate is saved, cached, and served to the client, establishing a secure connection.
Without strict optimizations, a sudden influx of traffic to invalid domains could trigger massive API degradation, or worse, cause your infrastructure to hit global ACME rate limits.
1. Implementing a High-Performance 'Ask' Endpoint
The ask permission check is your primary line of defense. If this endpoint is slow or poorly designed, your entire TLS handshake layer will suffer from severe latency. Follow these rules for production stability:
- Never query a relational database directly: Direct SQL queries inside the TLS handshake loop will quickly exhaust database connection pools. Instead, use an ultra-fast in-memory store like Redis or a distributed key-value store.
- Keep responses lightweight: The endpoint should return a simple
HTTP 200 OKif the domain is permitted, and aHTTP 400or404if it is not. No heavy JSON payloads are needed. - Implement local caching: If possible, let your 'Ask' microservice cache domain lookups locally in-memory for a few minutes to withstand brute-force attacks on specific custom domains.
2. Production-Ready Caddyfile Configuration
To support tens of thousands of domains smoothly, you must fine-tune global options within your Caddyfile. Below is an optimized configuration blueprint designed for high-concurrency, multi-tenant environments.
{
# Global Options
on_demand_tls {
ask [http://internal-api.local/v1/verify-domain](http://internal-api.local/v1/verify-domain)
interval 2m
burst 10
}
# Storage Optimization for Cluster Environments
storage redis {
host "redis-cluster.local"
port 6379
db 0
key_prefix "caddy_tls"
}
# Default ACME Configuration
acme_ca [https://acme-v02.api.letsencrypt.org/directory](https://acme-v02.api.letsencrypt.org/directory)
acme_ca_root /etc/ssl/certs/ca-certificates.crt
}
# Catch-all block for Multi-Tenant Custom Domains
:443 {
tls {
on_demand
}
# Reverse proxy to your internal application service
reverse_proxy http://app-backend-cluster {
header_up Host {http.request.host}
header_up X-Real-IP {http.request.remote}
header_up X-Forwarded-For {http.request.remote}
transport http {
dial_timeout 5s
keepalive_period 30s
}
}
}Key Configuration Insight: Theintervalandburstparameters restrict how many new certificates Caddy can request from ACME providers within a specific timeframe. Settingburst 10andinterval 2mprevents an attacker from pointing thousands of random domains to your IP and exhausting your Let's Encrypt quota.
3. Distributed Storage Infrastructure
By default, Caddy stores certificates in the local file system. In an enterprise environment with over 50,000 domains, relying on a single node's local storage is a critical anti-pattern. You need horizontal scalability (multiple Caddy instances behind a Layer 4 Load Balancer).
To achieve this, deploy a storage plugin like caddy-dns/redis or an S3-compatible backend. This guarantees that:
- Any Caddy instance in your cluster can fulfill a TLS handshake using a certificate requested by another node.
- Disk I/O bottlenecks are completely eliminated during peak hours.
- Backups and disaster recovery of certificate keys are centralized.
4. Optimizing Operating System and Network Limits
Handling tens of thousands of concurrent TLS handshakes requires optimization at the Linux kernel level. Ensure your host instances are tuned with the following sysctl parameters:
# Increase maximum open files (file descriptors) for high concurrency
fs.file-max = 2097152
# Optimize TCP window sizes and buffer allocations
net.core.somaxconn = 65535
net.ipv4.tcp_max_syn_backlog = 65535
net.ipv4.tcp_rmem = 4096 87380 16777216
net.ipv4.tcp_wmem = 4096 65536 16777216Additionally, remember to explicitly increase the systemd service limits for Caddy by setting LimitNOFILE=1048576 in your Caddy service definition file to prevent "too many open files" errors.
5. Mitigating Security Risks and DDOS Attacks
Exposing an On-Demand TLS endpoint opens up potential attack vectors. Implement these defensive security measures:
- Layer 4 Rate Limiting: Use Cloudflare (Advanced TCP Protection) or AWS Shield ahead of your Caddy cluster to filter out malformed TCP packets and volumetric DDoS attacks before they reach Caddy.
- Dual ACME Providers: Caddy natively supports fallback mechanisms. Configure both Let's Encrypt and ZeroSSL. If Let's Encrypt throws a 429 rate limit or goes down, Caddy will automatically switch to ZeroSSL to issue certificates seamlessly.
- Active Monitoring: Export Caddy metrics using the native Prometheus endpoint (
http://localhost:2019/metrics). Track the metricscaddy_tls_handshake_duration_secondsandcaddy_tls_issuance_failed_countto spot configuration errors instantly.
Conclusion
Scaling custom domains to 50,000+ targets doesn't require complex proprietary code or expensive enterprise load balancers. By combining Caddy Server v3's On-Demand TLS with an optimized Redis 'Ask' endpoint, distributed storage, and proper Linux kernel tuning, you can deliver a reliable, secure, and low-latency experience for your customers. Monitor your issuance success metrics closely, keep your verification layers fast, and let Caddy automate your public key infrastructure effortlessly.
