Optimizing Caddy Server for Enterprise: Scaling On-Demand TLS Across 50,000+ Custom Customer Domains
Introduction to Enterprise Multi-Tenancy and SSL Automation
For modern white-label Software-as-a-Service (SaaS) platforms, hosting providers, and multi-tenant applications, allowing clients to point their custom domains to your infrastructure is a critical business capability. However, managing SSL/TLS certificates manually for thousands of distinct domains introduces significant operational overhead and security risks.
Traditional web servers like Nginx or Apache require explicit configuration blocks for every domain or external cron jobs leveraging Certbot to request certificates. When managing a fleet of over 50,000 custom domains, this approach inevitably falls over due to disk I/O bottlenecks, high memory consumption, configuration reload latency, and Let's Encrypt rate limits.
Caddy Server addresses this challenge at scale through its native On-Demand TLS feature. Instead of pre-configuring thousands of hostnames, Caddy dynamically provisions, installs, and renews TLS certificates on-the-fly during the initial TLS handshake. This technical guide outlines the architecture and optimization strategies required to run Caddy Server in production for high-volume, enterprise-scale custom domain deployments.
---The Architecture of On-Demand TLS
To scale certificate management seamlessly, it is essential to understand how Caddy handles a connection when a client requests a site via a new custom domain for the first time. The automated lifecycle operates through the following pipeline:
- TCP Connection & TLS Handshake: A user navigates to
client-domain.com, which points to your infrastructure via a CNAME or A record. The browser initiates a TLS handshake with the Caddy Edge proxy, sending the hostname via Server Name Indication (SNI). - Local Cache Lookup: Caddy checks its internal memory cache and distributed storage to see if a valid certificate already exists for
client-domain.com. If found, it completes the handshake instantly. - The Permission Check (The 'Ask' Directive): If the certificate is missing, Caddy intercepts the handshake and performs an HTTP internal verification request against your central application database or API endpoint. This process confirms whether the custom domain is registered and authorized on your platform.
- Dynamic Issuance via ACME: Upon receiving a
200 OKresponse from your API, Caddy contacts an ACME provider (such as Let's Encrypt or ZeroSSL) via an asynchronous background routine, solves the HTTP-01 challenge, obtains the certificate, stores it in persistent storage, and caches it in memory. - Handshake Completion: The connection is secured, and the user is proxied to the backend upstream application.
Critical Security Note: Running On-Demand TLS without an automated permission check creates an extreme vulnerability. Malicious actors could point thousands of arbitrary domains to your server's IP address, forcing your instance to request infinite certificates, which exhausts Let's Encrypt rate limits and crashes your edge layer via resource exhaustion.---
Production-Ready Caddyfile Configuration
To handle a scale of 50,000+ domains efficiently, you must leverage global configurations that isolate storage, implement strict handshake verification via the modern permission module, and optimize downstream proxying. Below is an enterprise-grade production template:
{
# Global Options Block
admin 127.0.0.1:2019
# Distributed Storage Configuration for Cluster Coordination
storage redis {
host "redis-cluster.internal"
port 6379
db 0
key_prefix "caddy_tls"
}
# Modern On-Demand TLS Permission Verification
tls {
on_demand {
permission http {
endpoint https://api.yourcompany.internal/v1/validate-domain
timeout 5s
}
}
}
# Performance Logging
log {
output file /var/log/caddy/access.log {
roll_size 100mb
roll_keep 10
}
format json
}
}
# Dynamic Site Block Handling All Inbound Port 443 Traffic
:443 {
# Activate On-Demand TLS for this block
tls {
on_demand
}
# Security Header Enhancements
header {
X-XSS-Protection "1; mode=block"
X-Content-Type-Options "nosniff"
X-Frame-Options "DENY"
Referrer-Policy "strict-origin-when-cross-origin"
Strict-Transport-Security "max-age=31536000; includeSubDomains; preload"
}
# Optimized Reverse Proxy to Backend Microservices
reverse_proxy {
to https://upstream-load-balancer.internal:8443
transport http {
tls_server_name upstream-load-balancer.internal
keepalive 30s
dial_timeout 4s
}
}
}
---
Deep-Dive Optimizations for 50,000+ Domains
1. Distributed Storage and High-Availability Clustering
By default, Caddy writes certificates to the local file system. In an enterprise environment serving 50,000+ domains, a single server introduces a single point of failure (SPOF). You must run multiple Caddy instances behind a Layer 4 load balancer (such as AWS NLB or Cloudflare Magic Transit).
To enable the Caddy fleet to act as a cohesive cluster, implement a distributed storage plugin such as Redis, PostgreSQL, or Amazon S3. When Caddy Instance A requests an dynamic certificate, it writes the asset to the shared storage layer. If a subsequent request hits Caddy Instance B, it instantly retrieves the certificate from the shared cache, maintaining a fast time-to-first-byte (TTFB) and preventing redundant ACME renewal requests.
2. Optimizing the Validation API Endpoints
The HTTP permission check is executed during the live TLS handshake. Any latency introduced by your validation API directly delays the user's connection. Implement these optimizations on your validation microservice:
- High-Speed Caching: Your database should not be queried on every handshake. Store authorized custom domains in an in-memory database like Redis with aggressive caching layers.
- Fail-Safe Statuses: Ensure that your internal validation endpoint strictly returns a
200 OKstatus code for valid tenants and a403 Forbiddenfor unauthorized domains. - No Redirects: The validation endpoint must return a direct payload. Redirects will cause timeouts within Caddy's handshake sequence.
3. Mitigating ACME Rate Limits and Failovers
Let's Encrypt enforces a strict limit of 50 failed validation requests per account per hour. If customers point unpropagated or incorrect DNS records to your proxy, your instance could trigger these caps rapidly. To guarantee uptime, configure Multi-Issuer Fallback within your Caddy automation policies. If Let's Encrypt rate limits are reached or their API experiences downtime, Caddy will automatically fall back to ZeroSSL or another alternative provider seamlessly, ensuring uninterrupted operations.
---Monitoring, Observability, and Alerting
Managing an edge layer at this scale requires deep system visibility. Caddy emits structured JSON logs and provides a native Prometheus metrics endpoint. Ensure your operations stack monitors the following indicators:
| Metric Type | Target Parameter | Operational Goal |
|---|---|---|
| caddy_tls_handshake_duration_seconds | Handshake Latency | Must remain under 500ms even during cold-starts. |
| caddy_tls_certificates_count | Total Cached Certificates | Tracks cache sizing across the distributed memory layer. |
| http.log.error | Handshake Failures (JSON) | Alert on sudden spikes, indicating DNS misconfigurations or ACME blocks. |
Conclusion
Leveraging Caddy Server's On-Demand TLS eliminates the complexity of enterprise-scale SSL management. By transitioning from static file blocks to dynamic, API-driven handshake permissions backed by shared storage infrastructure, you can confidently support 50,000+ customer domains with maximum efficiency, zero downtime, and robust automated security protocols.
