Back to articles
Technology Insight

Building a High-Availability, Fault-Tolerant Private Reverse Proxy Cluster: Combining Caddy, Keepalived, and IP Anycast

June 4, 2026

Introduction to Enterprise-Grade Infrastructure Resilience

In modern enterprise architecture, the reverse proxy is the gateway to your infrastructure. It handles SSL/TLS termination, load balancing, traffic routing, and security filtering. However, placing a single reverse proxy instance at the edge introduces a critical vulnerability: a Single Point of Failure (SPOF). If that single instance goes down due to hardware failure, software crashes, or network congestion, your entire digital ecosystem becomes inaccessible.

To mitigate this risk, enterprise infrastructure teams must design systems that are inherently resilient. This blog post provides an engineering blueprint for constructing a highly available, fault-tolerant private reverse proxy cluster. By combining the modern efficiency of Caddy Server, the rapid failover capabilities of Keepalived, and the advanced routing precision of the IP Anycast protocol, you can achieve an automated, self-healing edge infrastructure that guarantees near-zero downtime.

The Core Architectural Pillars

Before diving into the implementation details, it is crucial to understand the distinct roles each component plays within this highly available topology. Rather than relying on a single monolithic solution, we distribute responsibilities across layers of the network stack to achieve optimal reliability.

1. Caddy Server: The Modern Edge Proxy

Traditionally, Nginx and HAProxy have dominated the reverse proxy landscape. However, Caddy Server has emerged as a formidable alternative for enterprise environments. Written in Go, Caddy offers memory safety, high performance, and an intuitive configuration syntax. Key advantages include:

  • Native Automatic HTTPS: Caddy manages Let's Encrypt or ZeroSSL certificates out of the box, handling renewal and management without external cron jobs or scripts. For private networks, it can orchestrate internal ACME servers seamlessly.
  • Dynamic Configuration: Through its robust JSON API, Caddy allows infrastructure teams to modify routing rules on the fly without restarting the process, eliminating configuration-induced micro-downtime.
  • Extensible Architecture: Its modular plugin system allows for easy integration with telemetry tools, custom authentication backends, and advanced logging frameworks.

2. Keepalived: High Availability Layer via VRRP

While Caddy ensures efficient traffic distribution to backend application servers, we still need a mechanism to protect the proxy layer itself. This is where Keepalived steps in. Keepalived implements the Virtual Router Redundancy Protocol (VRRP), a layer-4 protocol designed to provide standard failover functionality.

Keepalived groups multiple physical proxy servers into a single virtual router instance, assigned a Virtual IP (VIP). At any given moment, one node is elected as the Master, actively handling traffic bound for the VIP, while the remaining nodes stay in a Backup state. The Master node continuously broadcasts health heartbeats. If the Master fails to send these pulses within a predefined window, a Backup node instantly promotes itself, claiming the VIP and resuming traffic ingestion with minimal packet loss.

3. IP Anycast: The Scalable Routing Layer

While Keepalived is highly effective within a single local area network (LAN) or data center broadcast domain, it relies on layer-2 ARP switching, which does not scale across geographically distributed regions or complex routed networks. To transcend these boundaries, we introduce IP Anycast.

Anycast is a network addressing and routing technique where a single destination IP address is shared by multiple physical routing endpoints. Using routing protocols like OSPF (Open Shortest Path First) or BGP (Border Gateway Protocol), each proxy node in the cluster advertises the exact same IP address to the upstream network routers. The network infrastructure automatically handles the heavy lifting, routing client requests to the topologically "nearest" or healthiest proxy node. If a node goes offline, the upstream router automatically updates its routing tables, redirecting traffic to the next closest active node.

---

Designing the Cluster Topology

To implement this hybrid architecture successfully, we structure our environment into three distinct layers, ensuring that failures at any point are automatically isolated and mitigated.

Architectural Principle: True fault tolerance requires isolating state and health monitoring. By separating routing topology from application delivery, we ensure that individual software panics do not destabilize the underlying network layer.

The standard topology consists of multiple edge nodes distributed across separate racks or availability zones. Each node runs an identical stack: the Caddy daemon, the Keepalived daemon, and a routing daemon (such as FRRouting for OSPF/BGP advertisements). The configuration ensures that both localized node crashes (handled by Keepalived) and wider network partitions (handled by Anycast) are covered seamlessly.

---

Step-by-Step Implementation Guide

Phase 1: Deploying and Customizing Caddy Server

First, we establish identical Caddy configurations across all cluster nodes. In a corporate environment, utilizing Caddy's native internal ACME feature allows you to issue local certificates across your private infrastructure without public DNS challenges.

Create a production-grade Caddyfile on each node:

{
	# Global options for clustering
	storage file_system /var/lib/caddy
	admin localhost:2019
}

# Secure Reverse Proxy Configuration
app.private.enterprise.internal {
	tls internal
	
	reverse_proxy {
		to 10.0.10.21:8080 10.0.10.22:8080 10.0.10.23:8080
		health_uri /healthz
		health_interval 5s
		health_timeout 2s
		lb_policy round_robin
		lb_try_duration 5s
		lb_try_interval 250ms
	}
	
	log {
		output file /var/log/caddy/access.log {
			roll_size 100mb
			roll_keep 10
		}
	}
}

This configuration enforces round-robin load balancing across your backend microservices, complete with passive and active health checks. Ensure the Caddy service is enabled and running cleanly via your system's service manager.

Phase 2: Configuring Keepalived for Localized Failover

Next, we bind our nodes together using VRRP. Install Keepalived on your primary and secondary edge instances. Below is a sample configuration for the Primary Node (Master) located in /etc/keepalived/keepalived.conf:

vrrp_script check_caddy {
    script "/usr/bin/pgrep caddy"
    interval 2
    weight 2
}

vrrp_instance VI_1 {
    state MASTER
    interface eth0
    virtual_router_id 51
    priority 101
    advert_int 1
    
    authentication {
        auth_type PASS
        auth_pass Secr3tP@ssw0rd
    }
    
    virtual_ipaddress {
        10.0.50.100/24
    }
    
    track_script {
        check_caddy
    }
}

For the Backup Node, adjust the state to BACKUP and lower the priority to 100. The check_caddy script acts as an internal health check; if the Caddy process terminates, Keepalived dynamically drops the node's priority, forcing an immediate, graceful failover of the Virtual IP (10.0.50.100) to the secondary node.

Phase 3: Injecting IP Anycast for Data Center Resilience

To scale our cluster horizontally across routed boundaries, we integrate Anycast. Instead of mapping applications directly to individual host IPs, we configure our core enterprise routers to listen for BGP or OSPF route advertisements from our proxy nodes. This can be orchestrated using an open-source routing suite like FRRouting (FRR) running alongside Caddy.

When utilizing Anycast, every single proxy node advertises the same service loopback IP address (e.g., 192.168.100.1) to the top-of-rack (ToR) switch. The network engineering architecture operates under these strict rules:

  1. Equal-Cost Multi-Path (ECMP): The upstream switch distributes connections across all active nodes advertising the address, effectively multiplying your ingress throughput capacity.
  2. Deterministic Failover: If an entire server rack experiences a power event or hardware fault, its routing daemon dies, withdrawing the BGP path advertisement. The upstream core router instantly shifts traffic away from that path within milliseconds, avoiding black-holed requests.
---

Validation, Monitoring, and Failure Testing

Building a high-availability infrastructure is only half the battle; verifying its resilience under stress is paramount. To validate your new cluster, execute the following operational stress tests in a staging environment:

  • Service Failure Isolation: Manually stop the Caddy service on the Master VRRP node. Monitor your system logs to confirm that the Virtual IP immediately migrates to the backup node with zero dropped HTTP requests.
  • Network Partition Simulation: Disconnect the network interface of a primary Anycast node. Utilize continuous ping and curl loops to measure the failover convergence time of your network's OSPF/BGP convergence. It should gracefully stabilize in under a second.
  • Load Testing: Deploy an HTTP benchmarking tool (like wrk or Vegeta) against your Anycast VIP. Monitor CPU utilization across all Caddy nodes to ensure ECMP is accurately distributing traffic flows.

For long-term monitoring, integrate Caddy's native Prometheus metrics endpoint with a Grafana dashboard. Track indicators such as active connection counts, HTTP error rates (5xx responses), and Keepalived state transition logs to maintain deep observability over your cluster's operational health.

Conclusion

By blending the modern efficiency of Caddy Server, the rapid local failover mechanics of Keepalived, and the macroscopic routing intelligence of IP Anycast, you eliminate vulnerabilities at the network edge. This hybrid architectural design provides an industrial-grade private reverse proxy cluster capable of automatically surviving software failures, hardware crashes, and severe network partitions. Implementing these layers requires precision, but the resulting peace of mind and infrastructure resilience are well worth the investment.