Back to articles
Technology Insight

Building a High-Availability Etcd Cluster on Low-Spec Vultr VPS for Distributed Infrastructure Backup

June 4, 2026

Introduction: The Imperative of Distributed State Management

In modern cloud-native architectures, maintaining a reliable, highly available single source of truth is paramount. Etcd, a strongly consistent, distributed key-value store, serves as the backbone for Kubernetes and various distributed infrastructure platforms. It stores the critical configuration data, state, and metadata required to keep complex systems synchronized.

While enterprise deployments often allocate significant compute resources to state management, start-ups and independent developers frequently operate under tighter budget constraints. Fortunately, due to Etcd's efficient design, it is entirely feasible to build a robust, production-grade 3-node Etcd cluster using the lowest-spec Virtual Private Servers (VPS) from provider Vultr. This guide will walk you through the architectural principles, prerequisite setups, configuration steps, and validation techniques required to deploy a resilient cluster without breaking the bank.

1. Architecture and Budget Optimization Strategy

To achieve high availability and strict data consistency, distributed consensus engines rely on the Raft consensus algorithm. A core requirement of Raft is the formulation of a quorum. To survive the failure of N nodes, the cluster must consist of at least 2N + 1 nodes. Therefore, a 3-node cluster can tolerate the loss of exactly 1 node while maintaining full operational capability.

Why Low-Spec Vultr VPS?

Vultr's entry-tier Cloud Compute instances typically offer 1 vCPU, 1 GB RAM, and high-performance NVMe storage. Since Etcd's performance is heavily bound by disk write latency (I/O operations per second or IOPS) rather than raw CPU processing power, Vultr's standard NVMe storage pools make these budget-friendly instances highly capable of handling infrastructure state updates for small to medium workloads.

For this guide, we assume the deployment of three Vultr instances with the following example internal IP assignments:

  • etcd-node-01: 10.10.0.11 (Hostname: etcd-01)
  • etcd-node-02: 10.10.0.12 (Hostname: etcd-02)
  • etcd-node-03: 10.10.0.13 (Hostname: etcd-03)
Note: For optimal reliability, it is highly recommended to deploy these instances across different physical host nodes or availability zones within the same regional data center to minimize shared hardware risks.

2. Security Pre-requisites: Implementing Mutual TLS (mTLS)

Because Etcd manages sensitive structural definitions and configurations, allowing unencrypted communication poses a severe security hazard. We must implement Mutual TLS (mTLS) to guarantee that only authorized cluster nodes and clients can communicate with each other.

Using a tool like cfssl or standard openssl, you must generate a Certificate Authority (CA) and specific certificates for two distinct operational layers:

  1. Peer-to-Peer Certificates: Used for internal node communication and Raft replication.
  2. Client-to-Server Certificates: Used by external entities (like a Kubernetes API server or backup scripts) to query data.

Once generated, ensure the following keys and certificates are securely copied to the /etc/etcd/certs/ directory on all three nodes:

  • ca.pem (The root certificate)
  • server.pem & server-key.pem (Node-specific identification)
  • peer.pem & peer-key.pem (Internal routing identification)

3. Step-by-Step Installation and Initialization

The installation process should be executed across all three Vultr VPS nodes running a stable Linux distribution such as Ubuntu 24.04 LTS.

Step 3.1: Fetching and Compiling/Extracting Etcd Binaries

Avoid distribution package managers if they lag behind stable upstream releases. Download the official pre-compiled binaries directly from GitHub:

ETCD_VER=v3.5.15
wget [https://github.com/etcd-io/etcd/releases/download/$](https://github.com/etcd-io/etcd/releases/download/$){ETCD_VER}/etcd-${ETCD_VER}-linux-amd64.tar.gz
tar -xvf etcd-${ETCD_VER}-linux-amd64.tar.gz
sudo mv etcd-${ETCD_VER}-linux-amd64/etcd* /usr/local/bin/

Step 3.2: System User and Directory Structure Configuration

Running services as the root user violates the principle of least privilege. Create a dedicated system user and set precise directory permissions:

sudo useradd --system --no-create-home --shell /bin/false etcd
sudo mkdir -p /var/lib/etcd /etc/etcd/certs
sudo chown -R etcd:etcd /var/lib/etcd /etc/etcd

4. Crafting the Cluster Configuration

Instead of complex command-line arguments, we will leverage a structured YAML configuration file located at /etc/etcd/etcd.yml. Below is a detailed template for etcd-node-01. You will need to carefully adapt the IP addresses and node names for nodes 02 and 03.

name: 'etcd-01'
data-dir: '/var/lib/etcd'
listen-peer-urls: '[https://10.10.0.11:2380](https://10.10.0.11:2380)'
listen-client-urls: '[https://10.10.0.11:2379](https://10.10.0.11:2379),[https://127.0.0.1:2379](https://127.0.0.1:2379)'
initial-advertise-peer-urls: '[https://10.10.0.11:2380](https://10.10.0.11:2380)'
advertise-client-urls: '[https://10.10.0.11:2379](https://10.10.0.11:2379)'
initial-cluster: 'etcd-01=[https://10.10.0.11:2380](https://10.10.0.11:2380),etcd-02=[https://10.10.0.12:2380](https://10.10.0.12:2380),etcd-03=[https://10.10.0.13:2380](https://10.10.0.13:2380)'
initial-cluster-token: 'etcd-infra-cluster-token'
initial-cluster-state: 'new'
client-transport-security:
  cert-file: '/etc/etcd/certs/server.pem'
  key-file: '/etc/etcd/certs/server-key.pem'
  client-cert-auth: true
  trusted-ca-file: '/etc/etcd/certs/ca.pem'
peer-transport-security:
  cert-file: '/etc/etcd/certs/peer.pem'
  key-file: '/etc/etcd/certs/peer-key.pem'
  client-cert-auth: true
  trusted-ca-file: '/etc/etcd/certs/ca.pem'

Key Parameters Explained:

  • Port 2379: Dedicated to client requests and API queries.
  • Port 2380: Dedicated to internal peer-to-peer data synchronization and voting.
  • initial-cluster: String containing addresses of all members; critical for bootstrapping state correctly.

5. Configuring the Systemd Service Wrapper

To ensure that the Etcd service automatically boots up upon system restart and restarts automatically if it encounters an unhandled runtime exception, manage it via a systemd unit file at /etc/systemd/system/etcd.service:

[Unit]
Description=etcd distributed key-value store
Documentation=[https://github.com/etcd-io/etcd](https://github.com/etcd-io/etcd)
After=network.target

[Service]
Type=notify
User=etcd
ExecStart=/usr/local/bin/etcd --config-file=/etc/etcd/etcd.yml
Restart=always
RestartSec=5
LimitNOFILE=65536

[Install]
WantedBy=multi-user.target

Once saved, reload the systemd daemon, enable, and start the service simultaneously across all nodes:

sudo systemctl daemon-reload
sudo systemctl enable --now etcd

6. Health Verification and Performance Tuning

To ensure that your newly configured cluster has successfully achieved consensus, use the etcdctl command-line utility. You must supply your certificate parameters to authenticate correctly:

ETCDCTL_API=3 etcdctl \
  --endpoints=[https://10.10.0.11:2379](https://10.10.0.11:2379),[https://10.10.0.12:2379](https://10.10.0.12:2379),[https://10.10.0.13:2379](https://10.10.0.13:2379) \
  --cacert=/etc/etcd/certs/ca.pem \
  --cert=/etc/etcd/certs/server.pem \
  --key=/etc/etcd/certs/server-key.pem \
  endpoint health

If functioning normally, the output will display a success matrix confirming that all endpoints are operational and healthy. You can check the current cluster leader by replacing endpoint health with endpoint status --write-out=table.

Optimizations for Low-Spec Hardware

Because the lowest-spec Vultr instances share CPU resources, two critical adjustments can be added to the etcd.yml config file if you experience frequent heartbeats timeouts or performance warnings in your logs:

  1. Heartbeat Interval & Election Timeout: If the shared CPU introduces virtualization jitter, slightly increase the heartbeat interval to 250ms and the election timeout to 1250ms to prevent unnecessary leadership re-elections.
  2. Disk Priority: Ensure Etcd gets prioritized disk IO allocation using ionice -c2 -n0 within your service startup script to prevent log-writing delays.

Conclusion: A Resilient, Cost-Effective Foundation

Setting up a 3-node Etcd cluster on entry-level Vultr VPS instances demonstrates that building highly available, distributed infrastructure backups doesn't require complex or expensive cloud setups. By strictly applying mutual TLS security, tailoring the Raft election timeouts to accommodate lower-spec hardware, and leveraging SSD/NVMe speeds, you establish a resilient data state platform capable of underpinning robust cloud-native applications. Regular snapshotting routines should be established next to finalize your structural protection strategy.