Back to articles
Technology Insight

Scaling Caddy Server: How to Handle 20,000+ Customer Domains with On-Demand TLS on a VPS Cluster

May 29, 2026

Introduction: The Multi-Tenant Custom Domain Challenge

For modern Software-as-a-Service (SaaS) platforms, e-commerce builders, and website hosting providers, allowing users to point their custom domains to your infrastructure is a critical feature. However, managing SSL/TLS certificates for tens of thousands of external domains simultaneously introduces significant infrastructure overhead. Traditional web servers require manual configuration reloads, complex Let's Encrypt rate-limit navigation, and substantial memory allocation.

Enter Caddy Server. Thanks to its native, automated TLS management capabilities, Caddy has become the industry standard for handling multi-tenant routing. But what happens when you scale past 20,000+ customer domains pointing to a single VPS cluster? Without precise optimization, you risk high memory consumption, slow SSL handshakes, and potential certificate issuance failures. This comprehensive guide details how to optimize Caddy Server using On-Demand TLS to handle massive domain scaling smoothly and securely.

Understanding Caddy's On-Demand TLS Architecture

In a standard configuration, Caddy provisions certificates at startup for all domains explicitly listed in its configuration file. This approach breaks down when dealing with 20,000+ dynamic customer domains. Instead, we leverage On-Demand TLS.

With On-Demand TLS, Caddy does not obtain a certificate for a domain until a client initiates a TLS handshake for it. This defers certificate generation to the exact moment it is needed, saving immense storage space and CPU cycles. However, at a scale of 20,000+ domains, relying on out-of-the-box settings can lead to resource exhaustion and security vulnerabilities, such as Distributed Denial of Service (DDoS) attacks targeting your SSL generation endpoints.

Step 1: Implementing a Strict Ask Endpoint (Security First)

The most critical prerequisite for On-Demand TLS at scale is the ask mechanism. If you enable On-Demand TLS without an ask endpoint, malicious actors can point thousands of random domains to your server's IP address, forcing your server to request certificates for them. This will quickly exhaust your Let's Encrypt or ZeroSSL rate limits and crash your server.

The ask parameter configures Caddy to make an HTTP GET request to an internal backend API before attempting to issue a certificate. Caddy appends the domain name as a query parameter (e.g., [https://api.yourdomain.com/check-domain?domain=customer.com](https://api.yourdomain.com/check-domain?domain=customer.com)). If your API returns a 200 OK status code, Caddy proceeds with the TLS handshake. If it returns a 400 or 404, Caddy drops the connection immediately.

Architecture Best Practice: Ensure your internal ask endpoint is highly optimized. It should query a fast, cached database like Redis to validate domain ownership in less than 10 milliseconds to prevent latency during the TLS handshake.

Step 2: Externalizing Storage with a Distributed Backend

By default, Caddy stores its SSL certificates and metadata on the local file system. In a VPS cluster environment where multiple Caddy instances sit behind a load balancer, local storage creates a synchronization nightmare. If Instance A issues a certificate, Instance B will not have access to it, leading to redundant issuance requests and rate-limiting blocks.

To handle 20,000+ domains across a VPS cluster, you must externalize Caddy's storage layer. Caddy supports third-party storage plugins through custom builds via xcaddy. The most reliable options for cluster environments are:

  • Redis Storage (caddy-dns/redis): Ideal for ultra-fast, in-memory operations and low-latency certificate retrieval.
  • S3 Storage (ssm-storage or s3-storage): Highly durable and cost-effective for long-term certificate preservation.
  • PostgreSQL Storage: Excellent if your platform already relies on a robust relational database infrastructure.

By implementing a shared storage backend, any VPS node in your cluster can seamlessly retrieve and serve an existing certificate issued by a sibling node.

Step 3: Advanced Caddyfile Optimization for High-Performance Routing

To support massive traffic over thousands of concurrent domains, the global configuration block of your Caddyfile needs fine-tuning. Below is an optimized structural template designed for high-scale enterprise environments:

{
    # Global options
    on_demand_tls {
        ask [http://internal-api.local/v1/validate-domain](http://internal-api.local/v1/validate-domain)
        interval 2m
        burst 10
    }
    
    storage redis {
        host "redis-cluster.local"
        port 6379
        db 0
        timeout 5s
    }

    servers {
        timeouts {
            read_idle 30s
            read_header 10s
            write 30s
        }
    }
}

# Catch-all block for customer custom domains
:443 {
    tls {
        on_demand
    }

    # Proxy traffic to your main application cluster
    reverse_proxy http://app-backend-cluster {
        header_up Host {host}
        header_up X-Real-IP {remote_host}
        
        health_uri /healthz
        health_interval 10s
        health_timeout 5s
    }
}

In this configuration, the interval and burst parameters under on_demand_tls act as local rate limiters, ensuring that even if your ask endpoint allows the domains, Caddy will not overwhelm certificate authorities during a sudden traffic spike.

Step 4: Linux Kernel and OS-Level Tuning

When running a high-concurrency Caddy cluster on standard VPS instances, the underlying Linux kernel parameters often become the bottleneck before Caddy does. To prevent dropouts and "connection refused" errors, apply the following optimization settings via /etc/sysctl.conf:

  1. Increase File Descriptors: Each connection and certificate file requires a file descriptor. Set the system-wide limit by adding fs.file-max = 2097152.
  2. Optimize the IP Conntrack Table: Prevent connection tracking tables from overflowing: net.netfilter.nf_conntrack_max = 1048576.
  3. Adjust TCP Window and Buffer Sizes: Allocate adequate memory for busy sockets using net.ipv4.tcp_rmem and net.ipv4.tcp_wmem.
  4. Increase Maximum Backlog: Allow the kernel to queue more packets before processing: net.core.somaxconn = 65535.

Additionally, always ensure your Caddy systemd service file contains LimitNOFILE=65536 to explicitly grant the Caddy process authorization to utilize the expanded file descriptor limits.

Step 5: Monitoring, Metrics, and Maintenance

Operating a cluster handling 20,000+ domains without visibility is dangerous. Luckily, Caddy exposes native performance metrics in Prometheus format. By enabling the metrics endpoint globally, you can monitor vital statistics in real-time, including:

  • Active TLS handshakes per second.
  • Certificate expiration schedules and renewal successes.
  • HTTP cache hit ratios and response latencies.
  • Proxy backend availability state.

Pairing Prometheus with a Grafana dashboard allows operations teams to visually identify anomalous spikes in certificate generation requests, pinpointing potential security attacks or misconfigurations before they impact legitimate end-users.

Conclusion

Scaling Caddy Server to seamlessly route over 20,000 customer domains with On-Demand TLS is not only achievable on a lean VPS cluster, but it is also highly reliable when configured correctly. By implementing a zero-trust ask security verification layer, decoupling state using a distributed Redis or S3 storage backend, tuning the Linux network stack, and monitoring performance via Prometheus, you establish a resilient, self-healing routing architecture. This setup minimizes infrastructure costs while delivering an uncompromised, lightning-fast browsing experience for your entire tenant ecosystem.

Scaling Caddy Server: How to Handle 20,000+ Customer Domains with On-Demand TLS on a VPS Cluster | DPTCloud