Scaling Globally on a Budget: Building Multi-Region Anycast Layer 4 Load Balancing with Envoy and Keepalived
Introduction to Cost-Effective Global Scalability
In the modern digital landscape, ensuring high availability and low latency for global users is no longer a luxury reserved for enterprises with massive budgets. Traditionally, achieving multi-region redundancy required expensive proprietary hardware or premium cloud services that quickly drain the capital of growing businesses and startups.
However, by strategically combining open-source software with budget-friendly Virtual Private Servers (VPS) and Anycast-enabled networking providers, you can build a robust, enterprise-grade Layer 4 Load Balancing (L4LB) architecture. This post provides a comprehensive blueprint for configuring a multi-region Anycast L4LB infrastructure using Envoy Proxy and Keepalived on cheap VPS infrastructure.
Understanding the Architectural Components
Before diving into the configuration, it is essential to understand how the individual pieces of this architecture interact to create a seamless, resilient system.
1. Anycast Routing: The Global Traffic Director
Unlike Unicast, where a single IP address maps to a single physical machine, Anycast allows multiple geographically dispersed servers to share the exact same IP address. Routers on the internet use BGP (Border Gateway Protocol) to send traffic to the nearest healthy node topology-wise. If a region goes offline, the internet's routing fabric automatically redirects traffic to the next closest region, providing native failover and latency reduction.
2. Envoy Proxy: High-Performance Layer 4 Balancing
While frequently celebrated as a Layer 7 service mesh sidecar, Envoy is an incredibly performant and lightweight Layer 4 proxy. Operating at the transport layer (TCP/UDP), Envoy handles incoming connections with minimal CPU and memory overhead, making it ideal for low-spec, cheap VPS instances. It manages downstream connections and balances them across upstream backend application servers based on health checks.
3. Keepalived: Local High Availability
Anycast handles regional failover, but what happens if a single VPS inside a specific region fails? That is where Keepalived comes in. By utilizing the Virtual Router Redundancy Protocol (VRRP), Keepalived monitors the health of the Envoy instances within the same data center. It binds a local Virtual IP (VIP) to the active master node. If the master Envoy node fails, the backup node instantly takes over the local VIP, ensuring no single point of failure within the region.
System Architecture Overview
Our target architecture spans across multiple regions (e.g., US-East, EU-West, and Asia-Pacific). In each region, we deploy at least two cheap VPS instances acting as the load-balancing tier, positioned in front of our actual application backends.
- Anycast IP:
192.0.2.1(Announced globally across all regions via the VPS provider's network settings). - Region A (Primary Node): VPS running Envoy + Keepalived (Master).
- Region A (Secondary Node): VPS running Envoy + Keepalived (Backup).
- Upstream Backends: The actual web or application servers processing the workloads.
Step-by-Step Configuration Guide
Let us walk through the process of setting up a single region. This configuration should be replicated across your other targeted geographic regions.
Step 1: Network Layer Setup (Anycast & VIP)
To begin, you must select a VPS provider that supports BGP broadcasting or offers a native Anycast IP feature (such as Vultr, BuyVM, or specialized network overlays). Once you assign the Anycast IP to your account, you will configure Keepalived to manage a local Virtual IP that maps directly to your network interface.
Step 2: Configuring Keepalived for Internal Failover
Install Keepalived on both the master and backup VPS nodes within the region:
sudo apt-get update && sudo apt-get install -y keepalivedOn the Master VPS, configure /etc/keepalived/keepalived.conf:
vrrp_script chk_envoy {
script "pidof envoy"
interval 2
weight 2
}
vrrp_instance VI_1 {
state MASTER
interface eth0
virtual_router_id 51
priority 101
advert_int 1
authentication {
auth_type PASS
auth_pass SecretClusterPassword
}
virtual_ipaddress {
192.0.2.1/32
}
track_script {
chk_envoy
}}
On the Backup VPS, use an identical configuration, but change the state to BACKUP and reduce the priority to 100. This ensures that the Master node takes precedence as long as its Envoy process is running healthily.
Step 3: Deploying and Configuring Envoy Proxy for L4 Balancing
Next, install Envoy. Because we are optimizing for low-cost VPS instances, we want a strict Layer 4 TCP proxy configuration to maximize throughput while minimizing resource consumption. Create the following envoy.yaml configuration file:
static_resources:
listeners:
- name: l4_forward_listener
address:
socket_address:
address: 0.0.0.0
port_value: 443
filter_chains:
- filters:
- name: envoy.filters.network.tcp_proxy
typed_config:
"@type": [type.googleapis.com/envoy.extensions.filters.network.tcp_proxy.v3.TcpProxy](https://type.googleapis.com/envoy.extensions.filters.network.tcp_proxy.v3.TcpProxy)
stat_prefix: ingress_tcp
cluster: backend_application_cluster
clusters:
- name: backend_application_cluster
type: STRICT_DNS
lb_policy: ROUND_ROBIN
connect_timeout: 0.25s
load_assignment:
cluster_name: backend_application_cluster
endpoints:
- lb_endpoints:
- endpoint:
address:
socket_address:
address: backend1.local
port_value: 8080
- endpoint:
address:
socket_address:
address: backend2.local
port_value: 8080
health_checks:
- timeout: 1s
interval: 5s
unhealthy_threshold: 3
healthy_threshold: 2
tcp_health_check: {}
This Envoy setup listens on port 443 (HTTPS) and passes the raw TCP streams directly to the backend servers listed under the cluster endpoints using a fast, deterministic Round Robin algorithm. The embedded tcp_health_check guarantees that if an upstream backend node crashes, Envoy will immediately stop routing connections to it.
Operational Challenges and Mitigation Strategies
While this architecture is highly cost-effective, running a DIY global network infrastructure presents specific edge-cases that engineers must proactively manage:
- The "Flapping" Problem: If a VPS instance experiences intermittent network degradation, BGP routes may flap, causing traffic to rapidly jump between global regions. To counteract this, adjust Keepalived timers and implement BGP dampening settings on your provider portal.
- Asymmetric Routing: In Layer 4 Anycast environments, packets from the client to the server might take one path, while response packets take another. Because we are operating at Layer 4 without modifying the network layer header inappropriately, standard TCP sessions remain stable as long as the stateful translation occurs within the same regional cluster.
- Session Persistence: True stateful session persistence (like stickiness tied to an HTTP cookie) cannot be achieved effectively at a pure Layer 4 Anycast tier. If your application demands session persistence, you must handle it downstream at the application layer via shared Redis sessions, or transition Envoy to a Layer 7 proxy configuration at the expense of higher resource utilization.
Conclusion: Enterprise Reliability on a Startup Budget
By shifting the burden of regional failover to the network layer via Anycast, and securing local high availability using Keepalived and Envoy Proxy, you effectively remove the financial barrier to global scalability. This stack operates smoothly even on VPS instances with as little as 1GB of RAM, saving thousands of dollars monthly compared to proprietary enterprise load-balancing services. Implement this architecture today to give your global users the fast, uninterrupted digital experience they expect.
