Back to articles
Technology Insight

Architecting High-Availability Storage: Deploying a Fault-Tolerant Replicated Volume with GlusterFS Across Cloud VPS

June 6, 2026

Introduction to High-Availability Cloud Storage

In the contemporary digital landscape, data availability is the cornerstone of business continuity. As enterprises migrate critical workloads to the cloud, relying on a single virtual private server (VPS) introduces a dangerous single point of failure (SPOF). If the underlying hardware experiences an outage, or if network disruption isolates the instance, your applications grind to a halt, leading to costly downtime and potential data loss.

To mitigate this risk, modern infrastructure engineering demands fault-tolerant, distributed storage architectures. Among the open-source solutions capable of addressing this need, GlusterFS (Gluster File System) stands out as a powerful, scalable, and highly reliable software-defined storage platform. This technical guide provides a comprehensive, step-by-step blueprint for building an ultra-fault-tolerant, replicated storage volume distributed across two independent Cloud VPS instances using GlusterFS.

Understanding GlusterFS and Replicated Volumes

GlusterFS aggregates various storage servers (referred to as bricks) over network interfaces into a single, unified parallel network file system. It eliminates the need for a centralized metadata server, utilizing an elastic hashing algorithm instead. This architectural choice inherently removes performance bottlenecks and enhances system resilience.

For high-availability scenarios across two servers, we utilize a Replicated Volume. In this configuration, GlusterFS mirrors data across the bricks in the replica set. When a client writes data to the mount point, the file system synchronously duplicates the blocks across both nodes. This ensures that even if one Cloud VPS suffers a catastrophic failure, the remaining node continues serving data seamlessly without any operational disruption.

Critical Design Note: While a 2-node replication setup provides immediate data redundancy, it is susceptible to a condition known as split-brain if network connectivity between the two nodes breaks. To guarantee production-grade data integrity, it is highly recommended to introduce a third, low-resource node acting strictly as an Arbiter, or to configure strict quorum policies. For the purpose of this foundational architecture, we will focus on the core 2-node replication mechanism.

Prerequisites and Environment Setup

Before initiating the deployment, ensure that your environment meets the following baseline requirements:

  • Infrastructure: Two distinct Cloud VPS instances deployed in the same private network (VPC) to minimize latency.
  • Operating System: Ubuntu 22.04 LTS or Ubuntu 24.04 LTS installed on both nodes.
  • Network Privileges: Root or sudo administrative access on both servers.
  • Dedicated Storage: A secondary, unformatted disk partition (e.g., /dev/sdb) attached to each VPS. Avoid using the root partition for GlusterFS bricks.

Network and Hostname Configuration

For GlusterFS to communicate efficiently, both nodes must resolve each other via hostnames. Assign distinct hostnames to your servers and map them correctly.

Execute the following command to set hostnames on each respective server:

# On Node 1
sudo hostnamectl set-hostname node01.storage.local

# On Node 2
sudo hostnamectl set-hostname node02.storage.local

Next, edit the /etc/hosts file on both nodes to include their internal IP addresses:

10.0.0.11    node01.storage.local    node01
10.0.0.12    node02.storage.local    node02

Step 1: Preparing the Dedicated Storage Bricks

GlusterFS requires a dedicated directory structure backed by a robust file system. The industry standard recommendation for GlusterFS bricks is XFS due to its superior handling of extended attributes and large files.

  1. Format the secondary drive: On both nodes, format the raw disk (assuming it is mapped as /dev/sdb) to XFS:
    sudo mkfs.xfs -f /dev/sdb
  2. Create the mount point: Create a persistent directory where the raw storage will be attached:
    sudo mkdir -p /data/glusterfs/brick1
  3. Configure persistent mounting: Append the storage configuration to /etc/fstab to guarantee it mounts automatically upon system reboots:
    /dev/sdb /data/glusterfs/brick1 xfs defaults 0 0
  4. Mount the partition: Execute the mount command to initialize the storage space:
    sudo mount -a

Step 2: Installing and Initializing GlusterFS

With the physical storage layer ready, we proceed with installing the GlusterFS daemon on both Cloud VPS instances. To ensure maximum stability and access to the latest security patches, add the official GlusterFS upstream repository.

Run the following sequence of commands on both Node 01 and Node 02:

sudo apt update
sudo apt install -y software-properties-common
sudo add-apt-repository ppa:gluster/glusterfs-11
sudo apt update
sudo apt install -y glusterfs-server

Once installation finishes, enable and start the systemd service so that the storage cluster automatically initializes on system boot:

sudo systemctl enable glusterd
sudo systemctl start glusterd
sudo systemctl status glusterd

Step 3: Establishing the Trusted Storage Pool

At this stage, both servers are running independent instances of GlusterFS. We must connect them into a single, cohesive Trusted Storage Pool (TSP). This operation only needs to be executed from one of the nodes.

From Node 01, probe the second node to establish the cluster connection:

sudo gluster peer probe node02.storage.local

If the configuration is correct, you will receive a peer probe: success message. Verify the cluster membership and synchronization state by executing the following diagnostic command on either server:

sudo gluster peer status

The output should indicate that Node 02 is connected and operating within the cluster mesh.

Step 4: Creating and Starting the Replicated Volume

With the trusted pool successfully established, we can now define our fault-tolerant replicated volume. Within the mounted XFS directory, create a sub-directory specifically intended for the brick data to isolate it from system metadata:

# Run on BOTH nodes
sudo mkdir -p /data/glusterfs/brick1/gvol0

Now, execute the volume creation command from Node 01. We define a replica count of 2, specifying both bricks:

sudo gluster volume create replica-vol replica 2 node01.storage.local:/data/glusterfs/brick1/gvol0 node02.storage.local:/data/glusterfs/brick1/gvol0 force

After successfully provisioning the volume, initialize it to start receiving data streams:

sudo gluster volume start replica-vol

To inspect the layout, health parameters, and active brick connections of your new distributed array, check the comprehensive volume status:

sudo gluster volume info

Step 5: Client Mounting and High-Availability Testing

Your fault-tolerant replicated volume is now active and broadcasting. To utilize this high-availability storage pool, clients must mount it using the native GlusterFS client driver, which handles automatic client-side failover.

On your application server (or for testing purposes, on the nodes themselves), install the client packages:

sudo apt update && sudo apt install -y glusterfs-client

Create a directory dedicated to mounting the high-availability file system:

sudo mkdir -p /mnt/ha-storage

Mount the volume. Notice that you only specify one host in the mount command; the native client automatically fetches the cluster topology and tracks the secondary node for failover redundancy:

sudo mount -t glusterfs node01.storage.local:/replica-vol /mnt/ha-storage

Validating Fault Tolerance

To ensure your cluster behaves as intended during a hardware emergency, perform a manual failover validation test:

  1. Write a sample payload file into the mount directory: echo "Data Integrity Test" > /mnt/ha-storage/test.txt.
  2. Log into Node 02 and verify that the file has instantaneously synchronized to the underlying brick directory.
  3. Simulate a catastrophic hardware failure on Node 01 by stopping its network interface or shutting down the GlusterFS daemon: sudo systemctl stop glusterd.
  4. Return to your client mount point. Perform directory reads and write operations. The file system remains fully responsive, automatically routing traffic exclusively through Node 02.

Conclusion and Production Optimization

By implementing GlusterFS across two Cloud VPS instances, you have established a robust, fault-tolerant, and highly available replicated volume capable of keeping your services online during an unexpected server outage. Data is automatically written in parallel, shielding your application from localized hardware failures.

As you transition this architecture into a production environment, consider executing these advanced optimization and security configurations:

  • Network Security: Enforce strict firewall rules via ufw or cloud security groups, restricting access to GlusterFS communication ports (e.g., 24007, 24008, and brick ports) solely to trusted cluster members.
  • Performance Tuning: Tune network block sizes and read-ahead caches using gluster volume set parameters tailored to your specific application access patterns.
  • Monitoring & Alerting: Implement active monitoring infrastructure to track brick utilization and peer connectivity status, ensuring you are immediately alerted the moment an array degradation occurs.