Back to articles
Technology Insight

Building a High-Availability PostgreSQL Cluster: Automated Failover with Patroni and etcd on VPS

June 4, 2026

Introduction to High-Availability PostgreSQL

In today's digital economy, data availability is synonymous with business continuity. For enterprise applications relying on PostgreSQL, a standalone database server represents a critical single point of failure (SPOF). If the primary server crashes due to hardware failure, network partitioning, or resource exhaustion, businesses face immediate downtime and potential data corruption. To mitigate this risk, deploying a High-Availability (HA) PostgreSQL cluster with automated failover is paramount.

Achieving true high availability requires more than simple master-slave replication. It demands an intelligent orchestration layer capable of continuously monitoring database health, detecting failures, and safely promoting a standby replica without human intervention. This is where Patroni and etcd become indispensable components of the modern database infrastructure stack.

The Core Components: Patroni and etcd

Before diving into the configuration steps, it is essential to understand the architectural roles of the tools driving this high-availability solution:

  • Patroni: Developed by Zalando, Patroni is an open-source template that configures and manages PostgreSQL HA clusters. It acts as a controller daemon wrapped around the PostgreSQL instance, using a Distributed Consensus Store (DCS) to track and enforce the cluster state.
  • etcd: A strongly consistent, distributed key-value store developed by CoreOS. Written in Go and utilizing the Raft consensus algorithm, etcd serves as the DCS for Patroni. It stores the authoritative state of the cluster, manages leader election, and ensures all nodes agree on which PostgreSQL instance is the current primary.
Why etcd? Automated failover requires strict consistency to prevent "split-brain" scenarios—a catastrophic situation where two nodes simultaneously believe they are the primary, leading to data divergence. etcd prevents this by requiring a quorum consensus before committing state changes.

Architectural Blueprint for the Cluster

To establish a resilient and fault-tolerant environment, a minimum of three nodes is highly recommended. This ensures that the etcd cluster maintains a majority quorum if a single node fails. Below is the blueprint for our VPS-based deployment:

  1. Node 1 (Primary): Runs etcd, Patroni, and PostgreSQL (Primary).
  2. Node 2 (Replica): Runs etcd, Patroni, and PostgreSQL (Standby).
  3. Node 3 (Replica/Witness): Runs etcd, Patroni, and PostgreSQL (Standby).

All nodes must communicate over a secure, low-latency private network interface. Ensure that your VPS provider supports private networking and that firewalls are strictly configured to permit traffic only between these specific nodes.

Step-by-Step Implementation Guide

Step 1: Preparing the VPS Environment

Begin by updating the package repositories and installing the necessary prerequisites on all three nodes. For this guide, we assume an enterprise Linux distribution such as Ubuntu 22.04 LTS or newer.

Execute the following commands to install PostgreSQL and Python dependencies:

sudo apt-get update && sudo apt-get install -y postgresql-16 python3-pip python3-psycopg2

Stop the default PostgreSQL service immediately after installation, as Patroni will take full control of managing the PostgreSQL process lifecycle:

sudo systemctl stop postgresql
sudo systemctl disable postgresql

Step 2: Deploying and Configuring the etcd Cluster

Next, install etcd on all three nodes to create your Distributed Consensus Store:

sudo apt-get install -y etcd-server etcd-client

Modify the /etc/etcd/etcd.yml configuration file on each node. You must define the node's unique name, data directory, listen URLs, and the initial cluster members. A robust etcd configuration looks like this:

name: 'node-1'
data-dir: '/var/lib/etcd/default.etcd'
listen-peer-urls: '[http://10.0.0.1:2380](http://10.0.0.1:2380)'
listen-client-urls: '[http://10.0.0.1:2379](http://10.0.0.1:2379),[http://127.0.0.1:2379](http://127.0.0.1:2379)'
initial-advertise-peer-urls: '[http://10.0.0.1:2380](http://10.0.0.1:2380)'
advertise-client-urls: '[http://10.0.0.1:2379](http://10.0.0.1:2379)'
initial-cluster: 'node-1=[http://10.0.0.1:2380](http://10.0.0.1:2380),node-2=[http://10.0.0.2:2380](http://10.0.0.2:2380),node-3=[http://10.0.0.3:2380](http://10.0.0.3:2380)'
initial-cluster-token: 'etcd-pg-cluster-token'
initial-cluster-state: 'new'

Replace the private IP addresses (10.0.0.x) with the actual private IPs of your respective VPS instances. Once configured, restart the etcd service and verify the cluster health:

sudo systemctl restart etcd
etcdctl endpoint health

Step 3: Installing and Configuring Patroni

Install Patroni using Python's package manager to ensure you receive the latest stable release:

sudo pip3 install patroni[etcd3]

Create a dedicated configuration file for Patroni at /etc/patroni/patroni.yml. This file instructs Patroni on how to connect to etcd, bootstrap the PostgreSQL database, and handle replication. Below is an enterprise-grade configuration template:

scope: postgres-ha-cluster
namespace: /service
name: node-1

dcs:
  etcd3:
    hosts:
      - 10.0.0.1:2379
      - 10.0.0.2:2379
      - 10.0.0.3:2379
  ttl: 30
  loop_wait: 10
  retry_timeout: 10
  maximum_lag_on_failover: 1048576

postgresql:
  listen: 10.0.0.1:5432
  connect_address: 10.0.0.1:5432
  data_dir: /var/lib/postgresql/16/main
  bin_dir: /usr/lib/postgresql/16/bin
  pg_hba:
    - host replication replicator 10.0.0.0/24 md5
    - host all all 0.0.0.0/0 md5
  authentication:
    replication:
      username: replicator
      password: StrongReplicationPassword
    superuser:
      username: postgres
      password: StrongSuperuserPassword

tags:
  nofailover: false
  noloadbalance: false
  clonefrom: false
  nosync: false

Ensure that the data_dir directory is empty before starting Patroni for the first time on the primary node, as Patroni will initialize the database via initdb or clone it from the leader.

Step 4: Automating Daemon Execution with Systemd

To guarantee that Patroni starts automatically upon system boot, create a systemd service unit file at /etc/systemd/system/patroni.service:

[Unit]
Description=Patroni PostgreSQL Orchestrator
After=network.target etcd.service

[Service]
Type=simple
User=postgres
ExecStart=/usr/local/bin/patroni /etc/patroni/patroni.yml
ExecReload=/bin/kill -HUP $MAINPID
KillMode=process
TimeoutSec=30
Restart=on-failure

[Install]
WantedBy=multi-user.target

Enable and start the service across all nodes:

sudo systemctl daemon-reload
sudo systemctl enable patroni
sudo systemctl start patroni

Validation and Testing the Failover Mechanism

Once the services are active, you can monitor the real-time topology of your database cluster using the patronictl CLI tool:

patronictl -c /etc/patroni/patroni.yml list

The output will display the cluster name, the role of each node (Leader or Replica), its status, and the current replication lag. To validate the automated failover mechanism, simulate a catastrophic failure on the primary node by stopping its Patroni service or abruptly forcing a system reboot:

sudo systemctl stop patroni

Execute the patronictl list command on one of the remaining standby nodes. You will observe that within the defined ttl window, etcd detects the loss of the leader lease, strips the failed node of its leader status, and automatically promotes one of the healthy standby replicas to Primary. Applications utilizing an intelligent connection pooler like PgBouncer or a virtual IP (VIP) will transparently redirect write traffic to the new leader.

Conclusion and Best Practices

Implementing a Patroni and etcd framework transforms PostgreSQL into an incredibly resilient database infrastructure. However, a production setup requires ongoing adherence to operational best practices:

  • Network Security: Always use TLS encryption for etcd communications and Patroni REST APIs to protect cluster orchestration from unauthorized interception.
  • Connection Pooling: Deploy HAProxy or PgBouncer in front of your cluster to dynamically route application traffic to the active leader based on Patroni's health endpoints.
  • Regular DR Simulations: Automated failover mechanisms should be routinely tested in a staging environment to ensure timeouts and replication configurations align with your business Recovery Time Objectives (RTO).

By leveraging cloud-native consensus tools like etcd and robust orchestrators like Patroni, engineering teams can operate PostgreSQL on VPS environments with the reliability and uptime profiles typically reserved for premium managed database solutions.