Building a Highly Available Object Storage Cluster Across Different Regions Using SeaweedFS
Introduction to Modern Distributed Storage
In today's digital economy, data availability is synonymous with business continuity. For enterprises deploying applications across cloud environments, relying on a single data center or region introduces a critical point of failure. While traditional storage solutions like Ceph offer robust features, their architectural complexity and heavy resource consumption often make them impractical for small to medium-sized VPS deployments. Enter SeaweedFS—a highly scalable, fast, and distributed object storage system designed to handle billions of files efficiently while keeping resource utilization low.
This technical guide provides a step-by-step blueprint for configuring a Highly Available (HA) Object Storage Cluster using SeaweedFS across two VPS nodes located in entirely different geographical regions. By implementing this cross-region architecture, your business can achieve robust fault tolerance, ensuring that even if an entire cloud region suffers a catastrophic outage, your data remains accessible and intact.
Understanding the SeaweedFS Architecture
Before diving into the configuration, it is essential to understand how SeaweedFS handles data. Unlike traditional file systems that look up file metadata on every disk access, SeaweedFS separates metadata management from actual file storage. It consists of three core components:
- Master Servers: Manage the volume layout, allocate volume IDs, and coordinate cluster-wide operations.
- Volume Servers: Store the actual files in large pre-allocated blocks (volumes) and handle read/write requests directly.
- Filer (Optional but recommended): Provides a higher-level interface supporting standard file operations, S3-compatible APIs, and directory structures.
In a standard single-datacenter setup, a quorum requires at least three master nodes. However, when constrained to two nodes across different regions, we must leverage SeaweedFS's built-in replication configuration to ensure data is written to both geographic locations simultaneously, maintaining a strict data sovereignty and fault-tolerant profile.
Prerequisites and Network Preparation
To successfully follow this guide, you will need the following infrastructure components:
- Two VPS Instances: Located in distinct geographic regions (e.g., Node A in US-East, Node B in EU-West) running a modern Linux distribution such as Ubuntu 22.04 or 24.04 LTS.
- A Secure Network Tunnel: Because these nodes communicate over the public internet, configuring a secure WireGuard VPN or a mesh network like Tailscale is highly recommended to encrypt cluster traffic and eliminate firewall complexities.
- Public/Private Keys: Configured for secure SSH access between nodes.
Security Note: Never expose SeaweedFS internal communication ports (typically 9333, 8080, and 18080) directly to the public internet without strict firewall rules or an underlying VPN layer.
Step 1: Installing SeaweedFS on Both Nodes
First, we must install the SeaweedFS binary on both VPS instances. You can download the pre-compiled binaries from the official GitHub repository.
Execute the following commands on both Node A and Node B:
wget [https://github.com/seaweedfs/seaweedfs/releases/download/3.65/linux_amd64.tar.gz](https://github.com/seaweedfs/seaweedfs/releases/download/3.65/linux_amd64.tar.gz)
tar -xzvf linux_amd64.tar.gz
sudo mv weed /usr/local/bin/
weed versionEnsure that the output displays the correct version, confirming a successful installation across both hosts.
Step 2: Configuring the Master Servers with Cross-Region Awareness
To achieve high availability across two regions, we will run a Master server on both nodes. We must explicitly define the data center and rack locations so SeaweedFS understands the physical separation of the hardware.
On Node A (US-East), create a systemd service or run the following command:
weed master -ip=10.0.0.1 -port=9333 -mdir=/var/lib/seaweedfs/master -peers=10.0.0.1:9333,10.0.0.2:9333 -defaultReplication=001On Node B (EU-West), run the corresponding command to link the masters:
weed master -ip=10.0.0.2 -port=9333 -mdir=/var/lib/seaweedfs/master -peers=10.0.0.1:9333,10.0.0.2:9333 -defaultReplication=001Note on Replication: The flag -defaultReplication=001 is critical. In SeaweedFS terminology, 001 means the data will be replicated once on a different rack/data center. This ensures that every file written to the cluster exists in both the US-East and EU-West regions automatically.
Step 3: Launching the Volume Servers
With the Master servers communicating, we now launch the Volume servers that will handle the physical storage. We must pass the specific data center designation to each volume daemon.
On Node A (Assigning to Data Center 1):
weed volume -mserver=10.0.0.1:9333 -dir=/mnt/storage/data -max=100 -ip=10.0.0.1 -port=8080 -dataCenter=us-eastOn Node B (Assigning to Data Center 2):
weed volume -mserver=10.0.0.2:9333 -dir=/mnt/storage/data -max=100 -ip=10.0.0.2 -port=8080 -dataCenter=eu-westBy explicitly stating -dataCenter=us-east and -dataCenter=eu-west, the Master server can correctly enforce the 001 replication rule, placing one copy of the data on each distinct VPS node.
Step 4: Setting Up the S3-Compatible Filer Interface
To make this cluster useful for modern applications, we need an S3-compatible interface. The SeaweedFS Filer bridges the gap between raw object blocks and a structured API. Run the Filer on both nodes to achieve a completely decentralized access layer.
On Node A:
weed filer -mserver=10.0.0.1:9333 -ip=10.0.0.1 -port=8888On Node B:
weed filer -mserver=10.0.0.2:9333 -ip=10.0.0.2 -port=8888To expose the S3 API layer, simply attach the S3 daemon to the local filer on each node using weed s3 -filer=localhost:8888 -port=8333. Now, your applications can point to either node's IP on port 8333 using standard AWS S3 SDKs.
Step 5: Verifying Fault Tolerance and Failover
To confirm that your highly available object storage cluster is functioning as intended, perform a simulation failover test:
- Upload a test document via the S3 API on Node A.
- Verify that the file is instantly accessible and readable via Node B.
- Simulate a network partition or regional outage by stopping the SeaweedFS services on Node A:
sudo systemctl stop seaweedfs-master seaweedfs-volume. - Attempt to read or write data via Node B.
Because of the 001 replication strategy, Node B holds a complete cryptographic copy of the dataset and will seamlessly serve requests, maintaining uptime despite the total loss of the primary node.
Conclusion and Operational Best Practices
Configuring a cross-region, highly available storage cluster using SeaweedFS offers an elegant balance between performance, simplicity, and fault tolerance. By leveraging low-cost VPS instances across separate geographic boundaries, organizations can build resilient storage backends tailored for backups, media hosting, and cloud-native application states.
For production deployments, consider implementing an external load balancer or a Failover DNS system (like AWS Route 53 or Cloudflare) in front of your S3 gateways. This ensures that application traffic is automatically rerouted to the surviving region without manual intervention, achieving a truly automated disaster recovery state.
