Building a DIY Multi-Region Anycast Layer 4 Load Balancer Using Keepalived and Bird on Budget VPS
Introduction to High-Availability Architecture on a Budget
In the modern digital landscape, ensuring high availability and low latency for web applications is no longer a luxury reserved for enterprises with massive infrastructure budgets. Traditionally, achieving multi-region redundancy meant relying on costly cloud providers and proprietary global server load balancing (GSLB) solutions. However, by combining foundational networking principles with powerful open-source software, engineers can build a robust, self-hosted alternative.
This technical guide demonstrates how to design and implement a DIY Multi-Region Anycast Layer 4 (L4) Load Balancing architecture. By leveraging Keepalived for local failover and Bird for Border Gateway Protocol (BGP) routing across three low-cost Virtual Private Servers (VPS), you can establish a resilient, geographically distributed entry point for your traffic.
Understanding the Core Components
Before diving into the implementation details, it is crucial to understand the architectural pillars that make this setup both resilient and performant:
- Anycast Routing: A network addressing and routing method where a single destination IP address is shared by multiple physical routing endpoints. BGP routes incoming traffic to the "nearest" available node based on the network topology.
- Layer 4 Load Balancing: Operating at the transport layer (TCP/UDP), L4 load balancing routes traffic based on IP routing and port numbers without inspecting the application layer content, rendering it highly efficient and fast.
- Bird Internet Routing Daemon: An open-source dynamic routing daemon that handles BGP peering, allowing our VPS instances to announce our Anycast IP prefix to the provider's upstream routers.
- Keepalived: A routing software based on the Virtual Router Redundancy Protocol (VRRP), used here to monitor local service health and gracefully withdraw BGP announcements if a backend service fails.
Architectural Overview and Network Design
For this deployment, we utilize three low-cost VPS instances across separate geographic regions (e.g., US-East, EU-Central, and Asia-Pacific) provided by a vendor that supports custom BGP announcements and Anycast IP configuration (such as Vultr, BuyVM, or Hetzner Cloud).
Each VPS functions as an edge node announcing the same Anycast IP address. Under normal conditions, global users are automatically routed to the closest node via BGP path selection. Within each node, Keepalived continuously monitors the health of your application servers (the Layer 7 proxies or upstream backends).
Key Architectural Rule: If an entire VPS node fails, the upstream provider detects the loss of the BGP session and automatically routes traffic to the next closest node. If only the local application fails, Keepalived triggers Bird to stop announcing the Anycast IP locally, initiating a graceful failover.
Step-by-Step Implementation Guide
1. Network and Prerequisites Preparation
To begin, ensure you have acquired a portable IP prefix (typically a /24 for IPv4 or /48 for IPv6) or are utilizing an Anycast IP range provided directly by your VPS vendor. You will also need the Autonomous System Number (ASN) of your provider and your own private or public ASN depending on the deployment scope.
2. Installing and Configuring Bird for BGP Anycast
Install Bird on all three VPS instances using your package manager:
sudo apt update && sudo apt install bird2 -y
The core responsibility of Bird is to peer with the upstream router and announce the Anycast IP. Below is an optimized sample configuration for /etc/bird/bird.conf:
log syslog all;
router id 192.168.1.1; # The physical public IP of the VPS
protocol device {
scan time 10;
}
# Define the Anycast IP on a dummy interface
protocol direct {
ipv4;
interface "dummy0";
}
# BGP template for upstream peering
protocol bgp upstream_provider {
local as 65001; # Your Private/Public ASN
neighbor 192.0.2.1 as 64512; # Upstream Router IP and ASN
ipv4 {
export filter {
if net = 203.0.113.1/32 then accept; # Your Anycast IP
reject;
};
import all;
};
}
Ensure you create a persistent dummy0 network interface on your OS and assign the Anycast IP (e.g., 203.0.113.1) to it.
3. Configuring Keepalived for Local Health Checking
While BGP handles node-level failures, Keepalived handles application-level failures. Install Keepalived across your fleet:
sudo apt install keepalived -y
Configure Keepalived via /etc/keepalived/keepalived.conf to track the status of your local backend application (for example, Nginx or an L4 TCP stream daemon):
vrrp_script check_app {
script "/usr/bin/curl -s -o /dev/null -w '%{http_code}' [http://127.0.0.1:8080/health](http://127.0.0.1:8080/health) | grep 200"
interval 2
weight 2
}
vrrp_instance VI_1 {
state MASTER
interface eth0
virtual_router_id 51
priority 101
advert_int 1
track_script {
check_app
}
notify_master "/etc/keepalived/scripts/bgp_up.sh"
notify_fault "/etc/keepalived/scripts/bgp_down.sh"
}
4. Tying Keepalived and Bird Together
To create the automation link, write the notification scripts referenced in the Keepalived configuration. When the local application is healthy (notify_master), we instruct Bird to start or maintain the BGP announcement. When it fails (notify_fault), we immediately disable the session or apply a filter to withdraw the prefix.
Example for bgp_down.sh:
#!/bin/bash
birdc down upstream_provider
logger "Keepalived detected application failure. Disabling BGP Anycast announcement."
Example for bgp_up.sh:
#!/bin/bash
birdc up upstream_provider
logger "Keepalived detected application recovery. Enabling BGP Anycast announcement."
Make sure to apply executable permissions to these scripts using chmod +x.
Testing and Validating Failover Scenarios
A resilient system must be thoroughly validated. Execute the following scenarios to ensure your DIY load balancer behaves correctly:
- Geographic Routing Verification: Run a
tracerouteor use global ping tools from various parts of the world. Verify that traffic from Asia hits the Asian VPS, and traffic from Europe hits the European VPS. - Application Layer Failure: Manually stop your backend application service on one node. Observe the Keepalived logs, verify that the local Bird daemon tears down the BGP session, and confirm that external traffic immediately shifts to the remaining two healthy nodes without packet loss.
- Complete Node Outage: Simulate a hard crash by rebooting or shutting down one VPS instance. The upstream router should detect the BGP hold-time expiration and update global routing tables within seconds.
Operational Best Practices and Limitations
While this architecture offers immense cost savings and architectural freedom, it requires careful engineering oversight:
- BGP Convergence Time: Unlike proprietary clouds that offer near-instant routing shifts, internet BGP convergence can take anywhere from a few seconds to a minute depending on upstream providers. Adjust your BGP timers (such as hold-time and keepalive-time) carefully.
- Stateful Connection Challenges: Anycast routes packets dynamically. If an internet routing shift occurs mid-session, TCP connections traveling to Node A might suddenly land on Node B, resulting in a TCP RST (Reset). This configuration is therefore ideal for stateless workloads, HTTP/3 (UDP-based QUIC), or short-lived TCP API requests.
- Monitoring: Set up decentralized external monitoring (such as Prometheus with blackbox exporters) to continuously probe each individual node's physical IP address as well as the collective Anycast IP address.
Conclusion
By blending the routing capabilities of Bird with the precise local health checking of Keepalived, you can successfully bypass expensive commercial load balancers. Deploying this architecture on just three budget VPS instances grants you a multi-region, self-healing network footprint capable of handling millions of requests with high resilience. It proves that with a deep understanding of network layers, enterprise-grade availability can be built on an indie budget.
