Back to articles
Technology Insight

Building High-Availability VPS Infrastructure with Keepalived and Floating IPs

May 17, 2026

Introduction to High-Availability VPS Architecture

In today's digital landscape, service continuity is not merely a convenience but a fundamental business requirement. Downtime translates directly to lost revenue, damaged reputation, and frustrated users. For organizations leveraging Virtual Private Servers (VPS), achieving high availability (HA) presents unique challenges compared to traditional dedicated hardware. This guide explores a robust, cost-effective solution: implementing high-availability VPS infrastructure using Keepalived and Floating IPs.

High-availability architecture ensures that if one server fails, another automatically takes over with minimal to zero service interruption. The combination of Keepalived—a routing software written in C—and floating IPs (also called virtual IPs) creates a failover cluster. This setup is particularly valuable for web servers, database read replicas, load balancers, and any service where uptime is critical.

Understanding the Core Components

What is Keepalived?

Keepalived is an open-source project that provides two main functionalities: health checking and VRRP (Virtual Router Redundancy Protocol) implementation. For high-availability setups, we primarily utilize its VRRP capabilities. VRRP allows multiple servers to share a virtual IP address, with one server acting as the master and others as backups. The protocol manages the election of the master server and facilitates automatic failover when the master becomes unavailable.

The Role of Floating IPs

A floating IP is an IP address that can be instantly reassigned from one server to another within the same network. Unlike a server's primary static IP, which is tied to specific hardware or a virtual instance, a floating IP is controlled at the software level. In an HA configuration, your service's public-facing address is this floating IP. Keepalived manages its assignment:

  • During normal operation, the floating IP resides on the master server
  • If the master fails health checks, Keepalived reassigns the floating IP to a designated backup server
  • This transition typically occurs within seconds, making it transparent to end users

Architectural Design and Prerequisites

Before implementation, careful planning ensures a smooth deployment. A typical two-node HA setup requires:

  1. Two or More VPS Instances: Deployed in the same data center or region to ensure low-latency communication between nodes. They should have identical or very similar specifications.
  2. Private Network Connectivity: Most cloud providers offer private networking options. Nodes must communicate over a private network for VRRP advertisements and health checks, avoiding public network volatility.
  3. Identical Service Configuration: The application or service (e.g., Nginx, Apache, HAProxy) must be installed and configured identically on all nodes. Data synchronization (for stateful services) must be addressed separately using tools like rsync, DRBD, or distributed storage.
  4. Floating IP Resource: Acquire a floating IP from your VPS provider. Providers like DigitalOcean, Linode, Vultr, and OpenStack-based clouds offer this as a standard feature.

Step-by-Step Implementation Guide

Step 1: Network Configuration and Preparation

Begin by configuring the private network on each VPS node. Obtain the private IP addresses assigned by your provider. Ensure these IPs are reachable between nodes by disabling relevant firewall rules temporarily for testing. Update each server's hostname (e.g., web-01, web-02) and add entries to /etc/hosts for easy identification.

Step 2: Installing and Configuring Keepalived

Install Keepalived using your distribution's package manager. For Ubuntu/Debian systems, use apt install keepalived. For CentOS/RHEL systems, use yum install keepalived. The core configuration resides in /etc/keepalived/keepalived.conf. Below is a sample configuration for a master node:

vrrp_instance VI_1 {
    state MASTER
    interface eth1
    virtual_router_id 51
    priority 100
    advert_int 1
    authentication {
        auth_type PASS
        auth_pass secure_password_123
    }
    virtual_ipaddress {
        10.0.0.100/24 dev eth1
    }
}

For the backup node, the configuration differs slightly: set state BACKUP and a lower priority value (e.g., 90). The virtual_router_id must match across all nodes in the same cluster. The authentication block provides simple security for VRRP packets.

Step 3: Configuring Health Checks

Keepalived's true power emerges with custom health check scripts. These scripts determine if a node is healthy enough to hold the floating IP. A common check verifies if a web server is running. Create a script /etc/keepalived/check_nginx.sh:

#!/bin/bash
if systemctl is-active --quiet nginx; then
    exit 0
else
    exit 1
fi

Make it executable (chmod +x). Then, reference it in the Keepalived configuration within the vrrp_instance block:

track_script {
    chk_nginx
}
...
vrrp_script chk_nginx {
    script "/etc/keepalived/check_nginx.sh"
    interval 2
    weight -20
}

If the script exits with a non-zero code (failure), the node's priority is reduced by the weight value, potentially triggering a failover.

Step 4: Integrating the Floating IP

The virtual_ipaddress block in the configuration should contain the floating IP address provided by your VPS provider. Ensure it's on the correct network interface (often the public interface like eth0 or ens3). Some cloud providers require specific configuration to allow an interface to hold an additional IP; consult your provider's documentation.

Step 5: Starting the Service and Testing Failover

Enable and start Keepalived on both nodes: systemctl enable --now keepalived. Verify the status with systemctl status keepalived. The master node should now hold the floating IP. You can verify this by running ip addr show on the master and looking for the floating IP listed on the designated interface.

To test failover, simulate a failure on the master node. You can stop the monitored service (systemctl stop nginx) or shut down the Keepalived service itself (systemctl stop keepalived). Within seconds, observe the floating IP migrate to the backup node. Use the ip addr show command on the backup to confirm. Restoring the master should, depending on priority settings, cause the floating IP to return (a behavior known as preemption).

Advanced Configuration and Best Practices

Multi-Node Clusters and Priority Management

You can extend this setup beyond two nodes. Add additional backup nodes with successively lower priority values. Keepalived will manage the failover chain accordingly. For complex scenarios, you can use notify scripts that trigger custom actions during state transitions (master to backup, backup to master, etc.). These are useful for starting/stopping ancillary services or sending alerts.

Security Considerations

  • VRRP Authentication: Always use a strong password in the auth_pass directive to prevent unauthorized nodes from joining the VRRP group.
  • Firewall Rules: Configure your firewall (iptables, nftables, or cloud security groups) to allow VRRP protocol traffic (IP protocol number 112) only between your cluster nodes on the private network. Block it from all other sources.
  • Script Security: Health check scripts should be owned by root and have restricted permissions to prevent tampering.

Synchronizing Application State

Remember, Keepalived only manages the IP address failover. For stateful applications (like databases with writes, user sessions in memory), you need a separate mechanism to synchronize data. Options include:

  • Shared Storage: Using a networked filesystem like NFS or a cloud block storage volume that can be detached and reattached.
  • Application-Level Replication: Such as MySQL master-replica replication with read/write splitting.
  • Distributed Systems: Designing the application to be stateless, storing all state in an external, highly-available data store like Redis Cluster or a distributed database.

Common Troubleshooting Scenarios

Even with careful configuration, issues can arise. Here are common problems and their solutions:

Floating IP Not Appearing on Master: Check the interface name in the configuration matches the system's actual interface. Verify the floating IP is in the same subnet as the server's primary IP on that interface. Ensure no other host on the network uses the same IP.

Failover Not Occurring: Inspect Keepalived logs (journalctl -u keepalived). Verify health check scripts are executable and exit with the correct codes. Confirm firewall rules are not blocking VRRP advertisements (multicast traffic on 224.0.0.18) between nodes.

Split-Brain Scenario: This occurs when both nodes believe they are the master, often due to network partition preventing VRRP communication. Using a higher advert_int and implementing a more robust health check that can detect network isolation can mitigate this risk. Some deployments add a third, low-resource witness node to break ties.

Conclusion: Building Resilience into Your Infrastructure

Implementing a Keepalived and floating IP solution transforms your VPS infrastructure from a fragile, single-point-of-failure setup into a resilient, highly-available system. While the initial configuration requires careful attention to detail, the long-term benefits for business continuity are substantial. This approach provides an excellent balance of cost, control, and reliability, especially for small to medium-sized deployments where enterprise-grade load balancers may be overkill.

The principles learned here—health monitoring, automatic failover, and shared virtual addressing—form the foundation of more complex HA strategies. As your needs grow, you can layer in global load balancing, multi-region failover, and more sophisticated state management. By mastering Keepalived today, you equip yourself with essential skills for building robust, fault-tolerant systems in the cloud era.