Back to articles
Technology Insight

Architecting Automated Cross-Cloud Disaster Recovery: High Availability Between Hetzner and Vultr VPS

May 28, 2026

The Strategic Necessity of Cross-Cloud Disaster Recovery

In an era where digital presence is synonymous with business continuity, relying on a single infrastructure provider represents a significant point of failure. While providers like Hetzner and Vultr offer high-tier reliability, regional outages, routing issues, or administrative complications can still lead to catastrophic downtime. This guide explores the architecture of an automated Disaster Recovery (DR) system that bridges these two powerhouses, ensuring that if one fails, the other assumes the workload seamlessly.

Understanding the Multi-Cloud Architecture

The goal is to create a Warm Standby or Active-Passive configuration where your primary workloads reside on Hetzner (often chosen for its cost-to-performance ratio) and failover automatically to Vultr (prized for its global footprint and advanced networking features).

Key Components of the Cluster

  • Primary Node (Hetzner): The active production environment handling live traffic.
  • Secondary Node (Vultr): The standby environment synchronized with the primary.
  • Data Replication Layer: Ensures both sites have identical data states.
  • Health Monitoring & Orchestration: Tools like Keepalived, HAProxy, or custom heartbeat scripts.
  • Global Traffic Management: DNS-based failover or BGP-based Anycast (via tools like Cloudflare or Vultr Load Balancers).

Step 1: Establishing Secure Cross-Provider Networking

Since your servers are in different data centers and different companies, they cannot communicate via local private LANs by default. We must establish a secure tunnel. Using WireGuard is recommended for its performance and simplicity.

Building an encrypted tunnel between Hetzner and Vultr ensures that database replication and file synchronization occur over a private, secure interface rather than the public internet.
  1. Install WireGuard on both VPS instances.
  2. Generate peer keys and assign static internal IPs (e.g., 10.0.0.1 for Hetzner, 10.0.0.2 for Vultr).
  3. Configure firewall rules (UFW or Firewalld) to allow traffic specifically on the WireGuard port.

Step 2: Real-Time Data Synchronization

Data is the heart of the DR strategy. For a failover to be effective, the RPO (Recovery Point Objective) must be as close to zero as possible. We focus on two main data types: Files and Databases.

Database Replication (MySQL/PostgreSQL)

Standard Master-Slave replication is the most reliable approach here. The Hetzner node acts as the Master, streaming binary logs to the Vultr Slave. For higher complexity, Galera Cluster or Group Replication can be used, though these require significant tuning for high-latency cross-provider links.

File System Sync (lsyncd or GlusterFS)

For application files, static assets, or user uploads, lsyncd is an excellent choice. It monitors the file system for changes and uses rsync to push updates to the Vultr node instantly. For enterprise-grade requirements, GlusterFS provides a distributed file system, though it demands more overhead.

Step 3: Implementing Health Checks and Automated Failover

Detection is the most critical phase of Disaster Recovery. If the system fails to detect an outage, the backup is useless; if it detects a 'false positive,' you suffer unnecessary instability.

Using Keepalived and VRRP

While VRRP is typically used in local networks, we can adapt health-check scripts that monitor the primary service (e.g., Nginx or Docker). If the primary service on Hetzner returns a 5xx error or stops responding to pings via the WireGuard tunnel, the orchestration script triggers the failover mechanism.

The Failover Logic

When an outage is confirmed:

  • The Vultr node promotes its database from 'Read-Only' to 'Primary'.
  • Application services on the Vultr node are started or reconfigured to handle traffic.
  • The Global Traffic Manager is notified of the change.

Step 4: Managing Global Traffic Redirection

How do users find the new server? Since we cannot move an IP address between Hetzner and Vultr, we must use DNS Failover or a Global Load Balancer.

Cloudflare DNS API Integration

One of the most cost-effective methods is using a script that monitors server health. Upon failure, it uses the Cloudflare API to update the 'A' record of your domain from the Hetzner IP to the Vultr IP. Modern providers like Cloudflare have a TTL (Time to Live) short enough to propagate this change in under 60 seconds.

Vultr BGP and Anycast

For advanced users, Vultr allows you to announce your own IP space via BGP. While complex, this allows for near-instantaneous routing changes without waiting for DNS propagation, though it requires owning your own IP prefixes.

Step 5: Testing and Maintenance (The 'Chaos' Strategy)

A Disaster Recovery plan that hasn't been tested is merely a suggestion. It is vital to perform 'Game Day' tests quarterly.

  • Simulated Network Partition: Block the WireGuard tunnel to see if the Vultr node reacts correctly.
  • Hard Shutdown: Forcefully power off the Hetzner VPS to verify the RTO (Recovery Time Objective).
  • Data Integrity Audit: Periodically verify that the checksums of files on Vultr match those on Hetzner.

Conclusion: Investing in Peace of Mind

Configuring an automated cross-cloud cluster between Hetzner and Vultr is not just a technical exercise; it is an insurance policy for your business. By leveraging the low-cost infrastructure of Hetzner with the flexible networking of Vultr, you create a robust, geographically redundant environment that can withstand even major provider-wide incidents. Start small with basic data replication, and gradually automate your failover until you achieve a high-availability architecture that scales with your growth.

Architecting Automated Cross-Cloud Disaster Recovery: High Availability Between Hetzner and Vultr VPS | DPTCloud