Back to articles
Technology Insight

Building a High Availability System with 2 VPS: Automatic Failover for Under 5 Million VND/Month

May 17, 2026

Introduction: The Need for Affordable High Availability

In today's digital landscape, application downtime translates directly to lost revenue, damaged reputation, and frustrated users. While enterprise-grade high availability (HA) solutions often come with six-figure price tags, small and medium-sized businesses, startups, and independent developers require a more accessible path to resilience. The good news is that robust, automatic failover is no longer exclusive to large budgets. By strategically leveraging two Virtual Private Servers (VPS) and open-source software, you can build a system that automatically recovers from server failure, maintaining service continuity for your users. This guide details a practical, step-by-step architecture to achieve this for a total operational cost of under 5 million Vietnamese Dong (VND) per month.

Core Architecture: Active-Passive Failover

The foundation of our cost-effective HA system is the active-passive (or master-backup) model. In this setup, one VPS acts as the active or master node, handling all production traffic and requests. The second VPS operates as the passive or backup node, running in an idle or warm-standby state, continuously synchronizing data and monitoring the health of the active node.

The magic of automatic failover is orchestrated by a virtual IP address (VIP). This VIP is a floating IP address that is assigned to the active server. All client traffic and DNS records point to this VIP. A monitoring daemon (like Keepalived) runs on both servers, constantly checking the health of the active node. Should the active server become unresponsive due to hardware failure, network issues, or service crashes, the monitoring system on the passive node detects the failure and automatically reassigns the VIP to itself. Within seconds, the passive server becomes the new active server, and service resumes with minimal interruption.

Component Breakdown and Technology Stack

Our architecture relies on a few key open-source technologies, each serving a specific purpose in the HA chain.

1. Keepalived: The Failover Conductor

Keepalived is the heart of the failover mechanism. It implements the Virtual Router Redundancy Protocol (VRRP), which allows a group of servers to share a virtual IP. Its responsibilities include:

  • Health Checking: Executing custom scripts to verify the status of critical services (e.g., web server, database).
  • VIP Management: Advertising and owning the virtual IP on the active node.
  • State Transition: Orchestrating the graceful transition from master to backup and vice-versa.

2. Data Synchronization Layer

A failover is useless if the backup server lacks current data. The synchronization strategy depends on your application:

  • For Web Files: Use lsyncd or rsync with inotify for near-real-time file replication from active to passive.
  • For Databases: Implement native replication. For MySQL/MariaDB, set up master-slave replication. For PostgreSQL, use streaming replication combined with a tool like repmgr for failover management.
  • For Configuration: Use version control (Git) and synchronized deployment scripts to ensure both nodes have identical application and service configurations.

3. Shared Storage (Optional but Recommended)

For truly stateful applications, consider a low-cost, network-attached block storage solution offered by many cloud providers. This volume can be attached to the active server and provides a consistent data disk for both nodes, simplifying the data layer. During failover, the storage volume is detached from the failed node and reattached to the new active node.

Step-by-Step Implementation Guide

Let's walk through a concrete example of setting up HA for a LEMP (Linux, Nginx, MySQL, PHP) stack.

Phase 1: Provisioning and Base Setup

Provision two identical VPS instances from a provider like Vultr, DigitalOcean, or a local Vietnamese cloud service. Choose a plan with at least 1 vCPU and 2GB RAM, which typically costs between 1.5 to 2.5 million VND per instance. Ensure they are in the same data center region for low-latency communication. Perform initial hardening: update packages, configure a non-root user, and set up a firewall (UFW).

Phase 2: Installing and Configuring Keepalived

Install Keepalived on both servers: sudo apt install keepalived (on Ubuntu/Debian). The configuration file (/etc/keepalived/keepalived.conf) defines the VRRP instance and health checks.

Example Master Node Configuration Snippet:
vrrp_instance VI_1 {
state MASTER
interface eth0
virtual_router_id 51
priority 100
advert_int 1
authentication {
auth_type PASS
auth_pass YourSecurePassword123
}
virtual_ipaddress {
192.168.1.100/24 dev eth0
}
track_script {
chk_nginx_service
}
}

The backup node configuration is identical, except with state BACKUP and a lower priority (e.g., 90). The track_script references a health check script that verifies if Nginx is running. If this script fails on the master, its priority is reduced, triggering the backup to take over the VIP.

Phase 3: Configuring Data Replication

For web files, set up lsyncd on the master to watch /var/www/html and rsync changes to the backup server's IP. For MySQL, configure master-slave replication. On the master, enable binary logging and create a replication user. On the slave, point it to the master's IP and start the replication thread. Test by creating a database on the master and verifying it appears on the slave.

Phase 4: Testing the Failover

The most critical phase is validation. Simulate failures and observe the system's response:

  1. Service Failure: Stop Nginx on the master. The health check script should fail, causing Keepalived to lower the master's priority. Within 1-3 seconds, the backup should assume the VIP. Verify by pinging the VIP and accessing your website.
  2. Server Failure: Power off the master VPS entirely. The backup, after missing VRRP advertisements, will declare itself master and seize the VIP.
  3. Network Partition: Test scenarios where connectivity between nodes is lost to understand "split-brain" risks and how your configuration handles them (using VRRP priorities and authentication).

Cost Analysis and Budget Management

Let's break down the monthly costs to stay under the 5 million VND target, using approximate market rates.

  • VPS Instance (x2): 2 x 2,000,000 VND = 4,000,000 VND. This covers two decent 2GB RAM/1 vCPU instances.
  • Block Storage (Optional 50GB): ~500,000 VND. Adds robustness for database data.
  • Domain & DNS: ~200,000 VND (annual cost amortized monthly).

Total Estimated Cost: ~4,700,000 VND/month. This leaves a small buffer for traffic overages or a third, low-priority monitoring node. The primary cost-saving is the elimination of expensive proprietary load balancers and HA software licenses. Your investment is in setup time and expertise, not recurring license fees.

Monitoring, Maintenance, and Best Practices

Building the system is only the beginning. Sustained high availability requires vigilant operations.

  • Monitoring: Implement external monitoring (e.g., UptimeRobot) to alert you if the VIP becomes unreachable. Use internal tools like Prometheus Node Exporter and Grafana to track server metrics on both nodes.
  • Regular Failover Drills: Schedule monthly tests to ensure the failover process works reliably. Update and test your health check scripts.
  • Secure Communication: Use strong, unique passwords for VRRP authentication and database replication users. Consider setting up a private VLAN or VPN between your VPS instances if the provider supports it.
  • Documentation: Maintain clear, up-to-date runbooks detailing failover procedures, IP addresses, and service restoration steps for your team.

Conclusion: Resilience Within Reach

High availability is a fundamental requirement for modern digital services, but its implementation need not be prohibitively expensive. By combining commodity cloud infrastructure with powerful, battle-tested open-source software like Keepalived, you can architect a system that provides automatic failover, minimizes downtime, and protects your business continuity. The sub-5-million-VND monthly budget makes this level of resilience accessible to a wide range of organizations, empowering them to deliver reliable service to their customers. Start with the active-passive model outlined here, rigorously test your failover scenarios, and build the operational confidence that comes with knowing your application can withstand a single server's failure without impacting the user experience.