Scaling on a Budget: Multi-Region Anycast Layer 4 Load Balancing with Envoy and Keepalived
Introduction: The Multi-Region Availability Challenge
In modern cloud architecture, achieving high availability (HA) and low latency typically points engineering teams toward proprietary, expensive cloud load balancers. However, for organizations operating on tight budgets or utilizing budget Virtual Private Server (VPS) providers, these managed solutions are financially out of reach. This technical guide demonstrates how to architect a production-grade, Multi-Region Anycast Layer 4 (L4) Load Balancing solution using open-source tools: Envoy Proxy and Keepalived.
By leveraging the power of BGP-like routing semantics (simulated or natively supported by select budget providers) alongside robust health checking, you can deploy a infrastructure that automatically routes user traffic to the nearest healthy node, ensuring seamless failover across geographically distributed data centers.
Understanding the Architecture Components
Before diving into configuration, it is critical to understand the role each component plays within this architectural stack:
- Anycast Routing: A network addressing and routing method where a single destination IP address is shared by multiple routing topologies. Routers deliver packets to the closest destination based on routing metrics (hops, latency).
- Keepalived: A routing software based on the Virtual Router Redundancy Protocol (VRRP). In our setup, it monitors local service health and dynamically assigns or drops the shared VIP (Virtual IP) or interacts with the provider's BGP API to announce routing paths.
- Envoy Proxy: A high-performance, small-footprint L4 and L7 proxy designed for cloud-native environments. We utilize Envoy at Layer 4 (TCP/UDP) for upstream connection pooling, rapid health checking, and minimal overhead handling.
Key Insight: Operating at Layer 4 allows the load balancer to handle millions of concurrent connections per second with minimal CPU and memory usage, as it routes packets without inspecting the application-layer payload (HTTP/HTTPS).
Phase 1: Designing the Network and Topology
For a resilient multi-region deployment, we establish nodes in at least two distinct geographical locations (e.g., US-East, EU-Central, and Asia-Pacific). Each region features an identical network setup:
- Edge Anycast Router / VIP Gateway: Provided either via native BGP announcement supported by the VPS provider (such as Vultr, BuyVM, or Hetzner) or via an overlay tunnel network.
- Active/Passive Keepalived Pair: Inside each region, two cheap VPS instances run Keepalived to maintain local high availability.
- Envoy Proxy Instances: Co-located on the same nodes as Keepalived, acting as the high-throughput traffic distributors to actual backend application servers.
Phase 2: Configuring Envoy Proxy for Layer 4 Load Balancing
Envoy requires an explicit configuration file (envoy.yaml) to handle raw TCP streams. Below is an enterprise-grade configuration optimized for high concurrency, low timeouts, and aggressive health checking of upstream application servers.
static_resources:
listeners:
- name: ingress_l4_listener
address:
socket_address:
address: 0.0.0.0
port_value: 443
filter_chains:
- filters:
- name: envoy.filters.network.tcp_proxy
typed_config:
"@type": [type.googleapis.com/envoy.extensions.filters.network.tcp_proxy.v3.TcpProxy](https://type.googleapis.com/envoy.extensions.filters.network.tcp_proxy.v3.TcpProxy)
stat_prefix: ingress_tcp
cluster: backend_application_cluster
clusters:
- name: backend_application_cluster
connect_timeout: 0.25s
type: STRICT_DNS
lb_policy: ROUND_ROBIN
load_assignment:
cluster_name: backend_application_cluster
endpoints:
- lb_endpoints:
- endpoint:
address:
socket_address:
address: backend-app-01.local
port_value: 8443
- endpoint:
address:
socket_address:
address: backend-app-02.local
port_value: 8443
health_checks:
- timeout: 1s
interval: 2s
unhealthy_threshold: 2
healthy_threshold: 2
tcp_health_check: {}
This layout instructs Envoy to listen on port 443, accept raw TCP connections, and pass them directly to the backend_application_cluster using a Round Robin load-balancing strategy. The active TCP health check guarantees that dead application nodes are pruned out within 4 seconds.
Phase 3: High Availability and Anycast Simulation via Keepalived
To ensure traffic actually reaches Envoy, we employ Keepalived. In a true BGP Anycast VPS environment, Keepalived triggers an announcement script when changes occur. In a standard VIP configuration, it handles local failover. Here is the configuration file for the master node (/etc/keepalived/keepalived.conf):
vrrp_script check_envoy {
script "/usr/bin/pgrep envoy"
interval 2
weight 2
}
vrrp_instance VI_1 {
state MASTER
interface eth0
virtual_router_id 51
priority 101
advert_int 1
authentication {
auth_type PASS
auth_pass SecureVrrpToken
}
virtual_ipaddress {
192.0.2.100/32
}
track_script {
check_envoy
}
notify_master "/usr/local/bin/anycast_announce.sh up"
notify_backup "/usr/local/bin/anycast_announce.sh down"
notify_fault "/usr/local/bin/anycast_announce.sh down"
}
The anycast_announce.sh script is the glue that makes multi-region Anycast work on cost-effective infrastructure. When Keepalived acquires the master state, it calls your VPS provider's API or exabgp daemon to announce the IP route globally. If Envoy crashes, the script drops the announcement, forcing upstream Tier-1 telecom providers to re-route global user traffic to the nearest remaining healthy region.
Phase 4: Optimization, Multi-Region Failover, and Edge Cases
Running an infrastructure over raw, budget networks presents clear challenges, namely packet loss and BGP flap flapping. To optimize your stack, implement the following adjustments:
- Tune Linux Kernel Parameters: Modify
/etc/sysctl.confto handle massive connection tracking queues:net.ipv4.tcp_tw_reuse = 1net.core.somaxconn = 4096 - Dampening Flaps: Avoid rapid multi-region routing shifts by establishing a dampening interval in your health checking mechanisms so transient network blips do not trigger global BGP recalculations.
Conclusion
By coupling the blistering L4 processing speed of Envoy Proxy with the granular operational tracking of Keepalived, engineers can effectively bypass expensive enterprise edge solutions. Deploying this architecture on commodity VPS instances enables small businesses to maintain global, multi-region footprint, extreme fault tolerance, and minimal latency profiles at a fraction of standard cloud operating costs.
