Building a High-Availability PostgreSQL Cluster with Patroni and Etcd on 3 Budget Cloud Servers
Introduction: The Imperative of High Availability on a Budget
In modern enterprise architecture, data is the most critical asset. Downtime not only results in immediate financial loss but also erodes customer trust. For database management systems like PostgreSQL, ensuring continuous availability is paramount. Traditionally, achieving robust High Availability (HA) required expensive enterprise hardware or premium managed services. However, by combining open-source orchestration tools with strategic architectural design, businesses can deploy a resilient, self-healing database cluster using cost-effective infrastructure.
This comprehensive guide details the step-by-step implementation of a production-ready, highly available PostgreSQL cluster. By leveraging Patroni for template-driven clustering and Etcd as a Distributed Consensus Store (DCS), we will construct a robust architecture across three budget-friendly cloud servers. This setup guarantees automatic failover, prevents split-brain scenarios, and maintains strict data consistency at a fraction of traditional infrastructure costs.
The Core Architectural Components
To build a resilient cluster, it is essential to understand the roles of the specialized components involved in this architecture:
- PostgreSQL: The underlying object-relational database management system responsible for data persistence and transactional integrity.
- Patroni: An open-source, Python-based orchestrator customized for PostgreSQL. Patroni acts as a supervisor, dynamically managing the lifecycle of the database instance, monitoring health, and executing automated failovers.
- Etcd: A strongly consistent, distributed key-value store that serves as the Distributed Consensus Store (DCS). Patroni uses Etcd to maintain leader election states, store cluster configurations, and achieve consensus across nodes.
Why a Minimum of 3 Nodes is Required
A common mistake in database clustering is deploying a two-node system. In distributed systems, consensus requires a strict majority, calculated by the formula:
Majority = floor(N/2) + 1
In a 3-node cluster, the majority required is 2. If one node fails, the remaining two nodes can safely elect a new leader. In contrast, a 2-node cluster requires a majority of 2; if one node drops, the remaining single node cannot form a quorum, rendering automatic failover impossible. Thus, a minimum of three distinct nodes is non-negotiable to eliminate the risk of split-brain scenarios.
Network and Environment Planning
Before executing configuration commands, we must establish a clear network topography. For this guide, we assume three standard, budget-friendly Linux instances (e.g., Ubuntu 24.04 LTS) provisioned with private networking capabilities. Dedicated private IPs minimize latency and eliminate exposure to external network traffic.
We will utilize the following network mapping schema:
- Node 1 (Leader Candidate / Etcd Member 1): Hostname:
pg-node-01| Private IP:10.0.0.11 - Node 2 (Replica Candidate / Etcd Member 2): Hostname:
pg-node-02| Private IP:10.0.0.12 - Node 3 (Replica Candidate / Etcd Member 3): Hostname:
pg-node-03| Private IP:10.0.0.13
Ensure that all nodes can communicate securely via their private networks on the necessary ports: 2379/2380 for Etcd, 8008 for Patroni rest API, and 5432 for PostgreSQL replication and client connections.
Step 1: Installing Dependencies Across All Nodes
To maintain absolute consistency across the cluster, execute the initial setup packages systematically on all three nodes. Begin by updating the local package repositories and installing the official PostgreSQL engine alongside Python components required by Patroni.
sudo apt-get update && sudo apt-get upgrade -y
sudo apt-get install -y postgresql-16 postgresql-client-16 python3-pip python3-psycopg2 curlAfter installation, immediately stop and disable the default, unmanaged PostgreSQL system service. Patroni must maintain absolute control over the startup, shutdown, and configuration of the database instances.
sudo systemctl stop postgresql
sudo systemctl disable postgresqlStep 2: Provisioning the Etcd Distributed Consensus Cluster
Etcd acts as the source of truth for our cluster state. Install the Etcd package on all three nodes using the standard package manager:
sudo apt-get install -y etcd-server etcd-clientOnce installed, edit the main configuration file located at /etc/etcd/etcd.yml. Each node requires a specialized configuration reflecting its unique network identity. Below is a production-grade template example for Node 1 (pg-node-01):
name: 'pg-node-01'
data-dir: '/var/lib/etcd/pg-cluster.etcd'
listen-peer-urls: '[http://10.0.0.11:2380](http://10.0.0.11:2380)'
listen-client-urls: '[http://10.0.0.11:2379](http://10.0.0.11:2379),[http://127.0.0.1:2379](http://127.0.0.1:2379)'
initial-advertise-peer-urls: '[http://10.0.0.11:2380](http://10.0.0.11:2380)'
advertise-client-urls: '[http://10.0.0.11:2379](http://10.0.0.11:2379)'
initial-cluster: 'pg-node-01=[http://10.0.0.11:2380](http://10.0.0.11:2380),pg-node-02=[http://10.0.0.12:2380](http://10.0.0.12:2380),pg-node-03=[http://10.0.0.13:2380](http://10.0.0.13:2380)'
initial-cluster-token: 'etcd-pg-cluster-token'
initial-cluster-state: 'new'Note: Modify the 'name', 'listen-peer-urls', 'listen-client-urls', and 'advertise' variables on Node 2 and Node 3 to mirror their respective private IP addresses.
Restart and enable the Etcd daemon across all servers:
sudo systemctl restart etcd
sudo systemctl enable etcdValidate the health and operational integrity of the newly formed consensus store using the following diagnostic command:
etcdctl endpoint health --write-out=tableStep 3: Deploying and Configuring Patroni
With the distributed state engine online, install Patroni via the Python package manager to ensure access to the latest stable release binaries:
sudo pip3 install patroni[etcd3]Create a dedicated configuration directory and define the Patroni parameters in a unified YAML blueprint located at /etc/patroni/patroni.yml. Below is the configuration layout template for Node 1:
scope: postgres-ha-cluster
namespace: /service
name: pg-node-01
etcd3:
hosts:
- 10.0.0.11:2379
- 10.0.0.12:2379
- 10.0.0.13:2379
restapi:
listen: 10.0.0.11:8008
connect_address: 10.0.0.11:8008
bootstrap:
dcs:
ttl: 30
loop_wait: 10
retry_timeout: 10
maximum_lag_on_failover: 1048576
postgresql:
use_pg_rewind: true
use_slots: true
parameters:
shared_buffers: 258MB
wal_level: replica
max_wal_senders: 10
max_replication_slots: 10
hot_standby: "on"
initdb:
- encoding: UTF8
- data-checksums
pg_hba:
- host replication replicator 10.0.0.0/24 md5
- host all all 0.0.0.0/0 md5
postgresql:
listen: 10.0.0.11:5432
connect_address: 10.0.0.11:5432
data_dir: /var/lib/postgresql/16/main
bin_dir: /usr/lib/postgresql/16/bin
pgpass: /var/lib/postgresql/.pgpass
authentication:
replication:
username: replicator
password: StrongReplicationPasswordSecure
superuser:
username: postgres
password: StrongSuperuserPasswordSecureEnsure proper directory permissions are provisioned so that the postgres system user owns the configuration files:
sudo mkdir -p /etc/patroni
sudo chown -R postgres:postgres /etc/patroniReplicate this configuration across Node 2 and Node 3, carefully adjusting the name identifier, restapi.listen, restapi.connect_address, postgresql.listen, and postgresql.connect_address to target the respective node IPs.
Step 4: Launching the Cluster Control Plane
To manage Patroni systematically, construct a native systemd unit service file at /etc/systemd/system/patroni.service on every machine:
[Unit]
Description=Patroni Orchestrator for PostgreSQL HA
After=network.target etcd.service
[Service]
Type=simple
User=postgres
ExecStart=/usr/local/bin/patroni /etc/patroni/patroni.yml
ExecReload=/bin/kill -HUP $MAINPID
KillMode=process
TimeoutSec=30
Restart=no
[Install]
WantedBy=multi-user.targetReload the system daemon management layer, register the services to execute upon boot sequence, and start Patroni:
sudo systemctl daemon-reload
sudo systemctl enable patroni
sudo systemctl start patroniStep 5: Operational Assessment and Failover Verification
Patroni bundles a command-line tool utility named patronictl to analyze, test, and control cluster deployments. To verify the live topology status of the cluster, run the following verification syntax under the postgres user privileges:
patronictl -c /etc/patroni/patroni.yml listThe management layout output will demonstrate a clean structural map indicating one node elected explicitly as the Leader with a status of running, while the remaining sibling nodes show as Sync Standby or Replica, actively receiving WAL stream updates via native streaming replication.
Simulating an Unplanned Node Outage
To validate the automated failover resilience of our high-availability architecture, we can simulate an unexpected infrastructure failure on the leader node. Execute a hard stop of the Patroni service on pg-node-01:
sudo systemctl stop patroniImmediately execute the status list tracking utility from either remaining standby node. The remaining active nodes interact instantly with the Etcd lease layer, identify the leader's demise, and trigger an automated promotion sequence. In under ten seconds, one of the healthy standbys assumes the primary leader lock, seamlessly keeping the database layer open for client writes without manual administrator intervention.
Conclusion: Enterprise-Grade HA Reached Efficiently
By implementing Patroni and Etcd across three low-cost cloud servers, you have established a highly resilient, enterprise-grade High Availability PostgreSQL architecture. This decoupled design guarantees your data infrastructure can seamlessly handle physical or network dropouts with automated, sub-second failover execution, all while completely avoiding vendor lock-in or premium software licensing costs.
As a next step to maximize the production value of this setup, consider pairing this architecture with a dynamic connection pooling database proxy such as PgBouncer or a virtual routing layer like HAProxy. Integrating a proxy layer ensures that client applications are automatically and transparently routed to whichever node currently holds the primary leader state, completing your highly available data stack.
