Back to articles
Technology Insight

Cross-Provider High Availability: A Guide to Manual Failover IP Configuration with Keepalived

May 30, 2026

Introduction to Cross-Provider High Availability

In today's digital landscape, relying on a single Virtual Private Server (VPS) provider poses a significant risk to business continuity. Even the most reputable cloud infrastructure giants experience routing failures, datacenter outages, and hardware degradation. True high availability (HA) requires architectural redundancy that spans across entirely independent infrastructure providers.

While many cloud vendors offer native "Failover IP" or "Floating IP" solutions, these features are almost universally restricted to their own internal networks. You cannot natively move a DigitalOcean Floating IP to an AWS EC2 instance, nor can you route a Linode Reserved IP to a Vultr instance. To achieve true cross-provider resilience, engineers must design a hybrid mechanism. This technical guide explores how to leverage Keepalived and the Virtual Router Redundancy Protocol (VRRP) alongside custom API scripting to orchestrate a manual, software-defined Failover IP mechanism between two distinct VPS providers.

The Core Challenge of Multi-Provider Routing

Before diving into the configuration, it is critical to understand the underlying networking constraints. Keepalived fundamentally relies on VRRP (Virtual Router Redundancy Protocol), which operates at Layer 3/4 to broadcast heartbeats between nodes. In a standard local network or a single-provider private VPC, when the master node fails, the backup node claims the Virtual IP (VIP) via Gratuitous ARP (GARP) requests.

The Multi-Provider Roadblock: Because your two VPS instances reside on completely separate physical networks and autonomous systems (ASNs), VRRP broadcasts cannot pass between them over the public internet, and standard GARP requests will be ignored by the upstream ISPs.

To overcome this limitation, our architecture will employ a dual-layer strategy:

  • Internal Health & State Synchronization: Keepalived will communicate over a secure, encrypted tunnel (such as WireGuard or OpenVPN) established directly between the two VPS instances to monitor node health and handle state transitions.
  • External Routing Manipulation: We will bypass traditional Layer 2/3 hardware switching by writing custom Keepalived tracking scripts. When a failure is detected, these scripts execute API calls to a third-party DNS provider (e.g., Cloudflare) or a dynamic BGP routing service to re-point public traffic to the backup provider's infrastructure.

Architectural Blueprint

For the purposes of this guide, our high-availability architecture consists of the following components:

  1. Primary Node (Node A): A VPS hosted with Provider 1 (e.g., DigitalOcean), acting as the primary application server.
  2. Secondary Node (Node B): A VPS hosted with Provider 2 (e.g., Vultr), acting as the standby/backup application server.
  3. Private Overlay Tunnel: A secure WireGuard peer-to-peer connection bridging Node A and Node B to facilitate VRRP heartbeat communication.
  4. Dynamic Traffic Director: A public-facing entry point, managed via Cloudflare's API, which acts as our software-defined "Failover IP" by dynamically modifying DNS A records or routing rules upon node failure.

Step 1: Establishing the Secure Communication Tunnel

Because VRRP requires direct network visibility between nodes, we must first establish a private tunnel. WireGuard is selected due to its lightweight nature, exceptional performance, and ease of deployment.

On Node A (Primary), configure the WireGuard interface (/etc/wireguard/wg0.conf):

[Interface]
PrivateKey = 
Address = 10.0.0.1/24
ListenPort = 51820

[Peer]
PublicKey = 
AllowedIPs = 10.0.0.2/32
Endpoint = :51820
PersistentKeepalive = 25

On Node B (Backup), apply the corresponding configuration:

[Interface]
PrivateKey = 
Address = 10.0.0.2/24
ListenPort = 51820

[Peer]
PublicKey = 
AllowedIPs = 10.0.0.1/32
Endpoint = :51820
PersistentKeepalive = 25

Enable and start the WireGuard interfaces on both nodes using systemctl enable wg-quick@wg0 --now. Verify connectivity by executing a ping command from Node A to Node B's internal IP address (10.0.0.2).

Step 2: Installing and Configuring Keepalived

With our private network established, we can install Keepalived to manage the active/passive clustering states. Install the package on both machines using your package manager (e.g., apt install keepalived or yum install keepalived).

Primary Node Configuration (Node A)

Edit the file /etc/keepalived/keepalived.conf on Node A:

global_defs {
    router_id node_a
    enable_script_security
    script_user root
}

vrrp_script check_app {
    script "/usr/local/bin/check_app_health.sh"
    interval 2
    weight 2
}

vrrp_instance VI_1 {
    state MASTER
    interface wg0
    virtual_router_id 51
    priority 101
    advert_int 1

    authentication {
        auth_type PASS
        auth_pass Secr3tVrrpPass
    }

    track_script {
        check_app
    }

    notify_master "/usr/local/bin/failover_trigger.sh MASTER"
    notify_backup "/usr/local/bin/failover_trigger.sh BACKUP"
    notify_fault  "/usr/local/bin/failover_trigger.sh FAULT"
}

Backup Node Configuration (Node B)

Edit the file /etc/keepalived/keepalived.conf on Node B:

global_defs {
    router_id node_b
    enable_script_security
    script_user root
}

vrrp_script check_app {
    script "/usr/local/bin/check_app_health.sh"
    interval 2
    weight 2
}

vrrp_instance VI_1 {
    state BACKUP
    interface wg0
    virtual_router_id 51
    priority 100
    advert_int 1

    authentication {
        auth_type PASS
        auth_pass Secr3tVrrpPass
    }

    track_script {
        check_app
    }

    notify_master "/usr/local/bin/failover_trigger.sh MASTER"
    notify_backup "/usr/local/bin/failover_trigger.sh BACKUP"
    notify_fault  "/usr/local/bin/failover_trigger.sh FAULT"
}

Step 3: Creating Health Check and Automation Scripts

Keepalived uses tracking and notification scripts to assess local health and act on cluster state changes. First, create the local service health checker script at /usr/local/bin/check_app_health.sh. This ensures that even if the server is powered on, Keepalived triggers a failover if the specific application web server daemon crashes.

#!/bin/bash
# Verify if the local application or Nginx proxy is responding
curl --silent --fail http://localhost:80/health_check > /dev/null
exit $?

Next, we construct the failover execution script: /usr/local/bin/failover_trigger.sh. This shell script interacts with our DNS routing provider via API to re-map our pseudo-Failover IP to the active infrastructure node whenever Keepalived shifts its state.

#!/bin/bash
STATE=$1
CF_API_TOKEN="your_cloudflare_api_token"
ZONE_ID="your_zone_id"
RECORD_ID="your_dns_record_id"

NODE_A_IP="192.0.2.10"  # Public IP of Provider 1 Node
NODE_B_IP="198.51.100.20" # Public IP of Provider 2 Node

case $STATE in
    "MASTER")
        # Point DNS to Node A
        curl -X PUT "[https://api.cloudflare.com/client/v4/zones/$ZONE_ID/dns_records/$RECORD_ID](https://api.cloudflare.com/client/v4/zones/$ZONE_ID/dns_records/$RECORD_ID)" \
             -H "Authorization: Bearer $CF_API_TOKEN" \
             -H "Content-Type: application/json" \
             --data "{\"type\":\"A\",\"name\":\"app.example.com\",\"content\":\"$NODE_A_IP\",\"ttl\":60,\"proxied\":false}"
        ;;
    "BACKUP"|"FAULT")
        # Point DNS to Node B if this node isn't Master
        # To prevent race conditions, only execute update from the actual node becoming MASTER
        # or use conditional logic checking local hostname.
        if [ "$(hostname)" = "node_b" ] && [ "$STATE" = "MASTER" ]; then
            curl -X PUT "[https://api.cloudflare.com/client/v4/zones/$ZONE_ID/dns_records/$RECORD_ID](https://api.cloudflare.com/client/v4/zones/$ZONE_ID/dns_records/$RECORD_ID)" \
                 -H "Authorization: Bearer $CF_API_TOKEN" \
                 -H "Content-Type: application/json" \
                 --data "{\"type\":\"A\",\"name\":\"app.example.com\",\"content\":\"$NODE_B_IP\",\"ttl\":60,\"proxied\":false}"
        fi
        ;;
esac

Ensure both scripts are given executable privileges by running chmod +x /usr/local/bin/*.sh before starting the Keepalived daemon.

Step 4: Testing Failover and Resilience

With configuration complete, start Keepalived on both systems (systemctl start keepalived). Monitor system logs on Node A using tail -f /var/log/syslog | grep Keepalived. Node A should successfully claim the MASTER state.

To test the failover mechanism under real-world disaster conditions, execute one of the following testing scenarios:

  • Scenario A (Software Failure): Stop the local web service on Node A (e.g., systemctl stop nginx). The check_app_health.sh script will return a non-zero exit code, lowering Node A's priority and causing Node B to seamlessly transition to MASTER.
  • Scenario B (Infrastructure Power Failure): Simulate a catastrophic hypervisor crash by abruptly shutting down or stopping Node A via its VPS cloud control panel. Node B will stop receiving VRRP advertisement packets over the WireGuard tunnel and assume control within seconds.

Review your external traffic director or Cloudflare dashboard during these tests to verify that the target IP resolves appropriately to Node B without manual intervention.

Conclusion and Operational Best Practices

Configuring a software-defined, manual Failover IP solution using Keepalived across two distinct cloud providers breaks structural dependency on single-vendor networks. However, maintaining this setup long-term requires careful operational hygiene:

  • Data Synchronicity: A failover mechanism only provides value if the backup node contains up-to-date application data. Ensure block storage replication, database mirroring (such as MySQL master-slave replication), or distributed filesystems (like GlusterFS) run actively over your secure private tunnel.
  • State Machine Tuning: Set low Time-To-Live (TTL) parameters on your public DNS records (e.g., 60 seconds) or employ an Anycast API gateway to avoid client-side DNS caching issues during failover events.
  • Preventing Split-Brain Scenarios: If the WireGuard tunnel fails but both servers remain online, both might assume the MASTER state. Implement strict defensive conditions in your notification scripts to cross-verify the state of the peer node via alternative public networks before modifying external routing records.

By implementing these programmatic controls and decoupling your infrastructure layers, you ensure high availability that withstands even full-scale provider outages.

Cross-Provider High Availability: A Guide to Manual Failover IP Configuration with Keepalived | DPTCloud