Back to articles
Technology Insight

Cross-Cloud High Availability: Configuring Manual Failover IP with Keepalived Across Different VPS Providers

May 30, 2026

Introduction: The Necessity of Cross-Provider High Availability

In the modern digital economy, downtime translates directly to financial loss and reputational damage. While most Virtual Private Server (VPS) providers promise high availability (HA) within their own infrastructure, relying on a single cloud vendor introduces a critical single point of failure (SPOF). Data center outages, routing misconfigurations, or regional network blackouts can take down an entire cloud provider's network, leaving your business stranded.

To achieve true resilience, enterprise architectures must look toward cross-cloud redundancy. This technical blog post provides an in-depth guide on how to configure a manual Failover IP mechanism using Keepalived across two distinct VPS providers (for example, DigitalOcean and Linode/Akamai). By implementing this architecture, you ensure that if your primary cloud provider experiences a catastrophic failure, your traffic can be swiftly and seamlessly redirected to your secondary provider.

---

Understanding Keepalived and the Challenge of Multi-Provider VIPs

Keepalived is a robust routing software written in C, primarily known for its implementation of the Virtual Router Redundancy Protocol (VRRP). In a traditional, single-datacenter setup, Keepalived dynamically assigns a Shared or Virtual IP (VIP) to the primary node. If the primary node fails, the backup node detects the lack of VRRP advertisements and claims the VIP instantly.

However, when operating across two different VPS providers, standard VRRP cannot broadcast across the separate network boundaries. You cannot simply broadcast a Layer 2 VRRP packet from AWS to Google Cloud, nor can you easily share a local subnet IP. To overcome this limitation, we utilize Keepalived not for its automated Layer 2 VIP switching, but for its robust health-checking engine and its ability to trigger custom notification scripts upon state changes (Master/Backup/Fault). These scripts interact with the cloud providers' APIs or global DNS systems to update a Failover IP or routing policy manually and programmatically.

---

Prerequisites and Architecture Overview

Before diving into the configuration, ensure you have the following components ready:

  • Primary VPS (Node A): Hosted with Provider A (e.g., DigitalOcean), running Ubuntu 22.04/24.04 LTS.
  • Secondary VPS (Node B): Hosted with Provider B (e.g., Linode), running Ubuntu 22.04/24.04 LTS.
  • A Global Failover Mechanism: This can be a Reserved/Floating IP that supports cross-account routing (if using hybrid connect), or more commonly in multi-provider setups, a Managed DNS Service with an API (like Cloudflare, Route 53, or GoDaddy) where the dynamic 'Failover IP' concept is executed by rapidly updating the DNS A-record via API scripts.
  • API Tokens: Administrative access tokens for both VPS providers and your DNS provider to authorize script execution.
Architecture Note: In this setup, Keepalived nodes will communicate with each other over a secure, encrypted tunnel (such as WireGuard or OpenVPN) to allow VRRP heartbeats to safely traverse the public internet between the two providers.
---

Step 1: Establishing a Secure Cross-Provider Tunnel

Since VRRP communication requires direct IP connectivity, we must bridge the network gap between Provider A and Provider B. We will use WireGuard due to its lightweight nature and high performance.

Execute the following commands on both servers to install WireGuard:

sudo apt update && sudo apt install wireguard -y

Generate private and public keys on both nodes, and configure a simple point-to-point tunnel (e.g., 10.0.0.1/24 for Node A and 10.0.0.2/24 for Node B). Once the tunnel is active, verify that you can ping Node B from Node A via the internal WireGuard IP. This encrypted tunnel will serve as the dedicated medium for Keepalived’s VRRP heartbeat tracking.

---

Step 2: Installing and Configuring Keepalived

With secure communication established, install Keepalived on both systems:

sudo apt install keepalived -y

Configuring the Primary Node (Node A)

Create and edit the configuration file at /etc/keepalived/keepalived.conf on Node A:

vrrp_script check_app {
    script "/usr/local/bin/check_crypto_app.sh"
    interval 2
    weight 2
}

vrrp_instance VI_1 {
    state MASTER
    interface wg0
    virtual_router_id 51
    priority 101
    advert_int 1
    authentication {
        auth_type PASS
        auth_pass Secr3tKey
    }
    track_script {
        check_app
    }
    notify_master "/usr/local/bin/failover_trigger.sh MASTER"
    notify_backup "/usr/local/bin/failover_trigger.sh BACKUP"
    notify_fault "/usr/local/bin/failover_trigger.sh FAULT"
}

Configuring the Secondary Node (Node B)

Similarly, configure /etc/keepalived/keepalived.conf on Node B, ensuring the state is set to BACKUP and priority is lower:

vrrp_instance VI_1 {
    state BACKUP
    interface wg0
    virtual_router_id 51
    priority 100
    advert_int 1
    authentication {
        auth_type PASS
        auth_pass Secr3tKey
    }
    track_script {
        check_app
    }
    notify_master "/usr/local/bin/failover_trigger.sh MASTER"
    notify_backup "/usr/local/bin/failover_trigger.sh BACKUP"
    notify_fault "/usr/local/bin/failover_trigger.sh FAULT"
}
---

Step 3: Writing the Custom Failover and Health Check Scripts

The core magic of this multi-provider setup lies within the notification scripts. When Node A loses its connection or its local application fails, Keepalived triggers the transition scripts.

The Health Check Script (check_app.sh)

This script runs periodically to ensure your local web service or application is responding appropriately. Save this to /usr/local/bin/check_crypto_app.sh:

#!/bin/bash
# Check if local Nginx/Application is running
curl --silent --fail http://localhost:80/health || exit 1

Make sure to apply executable permissions: sudo chmod +x /usr/local/bin/check_crypto_app.sh.

The Failover Automation Script (failover_trigger.sh)

This script executes whenever the node transitions state. If a node becomes MASTER, it calls your DNS provider or BGP/Anycast API to point the public Failover IP traffic to its own public IP address.

#!/bin/bash
STATE=$1
NOW=$(date +"%Y-%m-%d %H:%M:%S")

case $STATE in
    "MASTER")
        echo "$NOW [Keepalived] Transitioned to MASTER. Updating Failover IP via API..." >> /var/log/keepalived-failover.log
        # Example: Call Cloudflare API to update A-Record to this server's public IP
        /usr/local/bin/update_dns_to_self.sh
        ;;
    "BACKUP"|"FAULT")
        echo "$NOW [Keepalived] Transitioned to $STATE. Standing down." >> /var/log/keepalived-failover.log
        ;;
    *)
        echo "$NOW [Keepalived] Unknown state: $STATE" >> /var/log/keepalived-failover.log
        exit 1
        ;;
esac

Remember to make this script executable as well: sudo chmod +x /usr/local/bin/failover_trigger.sh.

---

Step 4: Testing and Validating Your Cross-Cloud Failover

To guarantee your multi-cloud high-availability architecture works smoothly under stress, you must run rigorous validation tests.

  1. Start Keepalived: Run sudo systemctl start keepalived on both nodes. Check the system logs via tail -f /var/log/syslog to confirm that Node A claims the MASTER state and Node B settles into the BACKUP state.
  2. Simulate Local Application Failure: Stop your application service on Node A (e.g., sudo systemctl stop nginx). Within seconds, the health check script will fail, causing Node A to enter FAULT status. Node B will detect the absence of heartbeats, transition to MASTER, and execute its script to remap the Failover IP target to itself.
  3. Simulate Network Split/Provider Outage: Block the WireGuard interface traffic on one end using iptables. Node B will automatically promote itself to MASTER to protect uptime, while your API integration reroutes global traffic to Node B.
---

Conclusion and Production Recommendations

Configuring a manual Failover IP mechanism using Keepalived across disparate VPS vendors offers unprecedented infrastructure resilience. By abstracting the high-availability layer above a single provider's network boundary, your application becomes immune to localized vendor disasters.

For enterprise deployment, consider these final best practices:

  • Minimize TTL: Set the Time-To-Live (TTL) of your managed failover DNS records to the absolute minimum (e.g., 30 to 60 seconds) to ensure client-side caching clears rapidly during a failover event.
  • Data Replications: Ensure that your application data, database state, or user uploads are continuously synchronized between Provider A and Provider B in real time using tools like GlusterFS, rsync, or database clustering (e.g., Galera, PostgreSQL replication). High availability for network routing is only half the battle; data integrity is the other.
Cross-Cloud High Availability: Configuring Manual Failover IP with Keepalived Across Different VPS Providers | DPTCloud