Back to articles
Technology Insight

Building a Resilient Infrastructure: Configuring a High-Availability PostgreSQL Cluster with Patroni and Etcd on Cloud Servers

May 29, 2026

Introduction to Database High Availability

In the contemporary digital landscape, data is the lifeblood of enterprise operations. Whether powering an e-commerce platform, a financial transaction system, or a critical SaaS application, database uptime is non-negotiable. System failures, network partitions, and hardware degradations are not a matter of 'if,' but 'when.' Therefore, implementing a robust High Availability (HA) strategy for your database tier is a foundational requirement for modern infrastructure engineering.

While PostgreSQL is renowned for its reliability and advanced feature set, its native streaming replication does not inherently provide automatic failover. If the primary node encounters a critical error, human intervention is traditionally required to promote a standby node. This manual process introduces latency, increases the Recovery Time Objective (RTO), and elevates the risk of human error. To mitigate these challenges, enterprise architects turn to automated orchestration tools. This technical guide explores how to build a highly available, self-healing PostgreSQL cluster utilizing Patroni as the template-driven orchestrator and Etcd as the Distributed Consensus Store (DCS) on Cloud Server infrastructure.

Architectural Components: Patroni and Etcd

Before diving into the configuration steps, it is essential to understand the architectural topology and the specific roles each component plays within the ecosystem. A resilient HA cluster typically requires a minimum of three nodes to establish a quorum and prevent split-brain scenarios.

1. PostgreSQL Streaming Replication

PostgreSQL utilizes write-ahead logging (WAL) to record changes before they are written to data files. In a typical HA setup, asynchronous or synchronous streaming replication is employed to transmit these WAL records from the primary (read-write) node to one or more standby (read-only) nodes. This ensures data redundancy across the infrastructure.

2. Etcd: The Distributed Consensus Store

Etcd is a strongly consistent, distributed key-value store that serves as the single source of truth for the cluster. It implements the Raft consensus algorithm to manage cluster state, coordinate leader election, and store configuration parameters. By utilizing Etcd, the cluster maintains a strict consensus on which PostgreSQL node currently holds the primary lock, effectively neutralizing the risk of two nodes acting as primary simultaneously.

3. Patroni: The Orchestration Daemon

Patroni is an open-source clustering solution written in Python. It acts as a supervisor daemon that wraps around the PostgreSQL process. Patroni continuously monitors the local PostgreSQL instance, communicates health metrics to Etcd, and dynamically manages replication linkages. If the primary node becomes unresponsive, Patroni detects the loss of the leader key in Etcd and automatically orchestrates the promotion of the most up-to-date standby node.

Prerequisites and Environment Topology

To implement this architecture on Cloud Servers, we will deploy a standard three-node topology. Each cloud server will host an instance of Etcd, Patroni, and PostgreSQL to ensure complete decentralization.

  • Node 1: IP 10.0.0.11 (Hostname: pg-node-01)
  • Node 2: IP 10.0.0.12 (Hostname: pg-node-02)
  • Node 3: IP 10.0.0.13 (Hostname: pg-node-03)
Security Note: Ensure that your Cloud Provider's Security Groups or local firewalls (e.g., UFW/iptables) allow internal communication on ports 2379/2380 (Etcd), 8008 (Patroni REST API), and 5432 (PostgreSQL) exclusively within the private network boundary.

Step-by-Step Implementation Guide

Step 1: Installing Dependencies and Etcd

First, update your package repositories and install Etcd on all three cloud nodes. It is highly recommended to use enterprise-grade Linux distributions such as Ubuntu Server or Rocky Linux.

sudo apt-get update
sudo apt-get install -y etcd-server etcd-client python3-pip python3-psycopg2

Once installed, edit the /etc/etcd/etcd.yml configuration file on each node. The configuration must point to the local node's private IP while referencing the peer IPs of the other cluster members. Below is an abstracted example for pg-node-01:

name: 'pg-node-01'
data-dir: '/var/lib/etcd/pg-cluster.etcd'
listen-peer-urls: '[http://10.0.0.11:2380](http://10.0.0.11:2380)'
listen-client-urls: '[http://10.0.0.11:2379](http://10.0.0.11:2379),[http://127.0.0.17:2379](http://127.0.0.17:2379)'
initial-advertise-peer-urls: '[http://10.0.0.11:2380](http://10.0.0.11:2380)'
initial-cluster: 'pg-node-01=[http://10.0.0.11:2380](http://10.0.0.11:2380),pg-node-02=[http://10.0.0.12:2380](http://10.0.0.12:2380),pg-node-03=[http://10.0.0.13:2380](http://10.0.0.13:2380)'
initial-cluster-token: 'etcd-pg-cluster-token'
initial-cluster-state: 'new'
advertise-client-urls: '[http://10.0.0.11:2379](http://10.0.0.11:2379)'

Restart and enable the Etcd service across all instances:

sudo systemctl restart etcd
sudo systemctl enable etcd

Verify the health of the Etcd cluster using the command-line utility:

etcdctl endpoint health --cluster

Step 2: Installing PostgreSQL and Patroni

Install the official PostgreSQL binaries. Crucially, do not initialize a default database cluster via the package manager, as Patroni will handle the initial bootstrapping and directory creation.

sudo apt-get install -y postgresql-16
sudo systemctl stop postgresql
sudo systemctl disable postgresql
sudo pip3 install patroni[etcd3]

Step 3: Configuring Patroni

Create a dedicated configuration file for Patroni at /etc/patroni/patroni.yml. This file defines how Patroni interacts with Etcd and dictates the internal configuration of PostgreSQL. Ensure the postgres user owns this file.

scope: postgres-ha-cluster
namespace: /service
name: pg-node-01 # Change accordingly for node 02 and 03

etcd3:
  hosts:
    - 10.0.0.11:2379
    - 10.0.0.12:2379
    - 10.0.0.13:2379

restapi:
  listen: 10.0.0.11:8008
  connect_address: 10.0.0.11:8008

bootstrap:
  dcs:
    ttl: 30
    loop_wait: 10
    retry_timeout: 10
    maximum_lag_on_failover: 1048576
    postgresql:
      use_pg_rewind: true
      use_slots: true
      parameters:
        max_connections: 100
        wal_level: replica
        hot_standby: "on"
        max_wal_senders: 10
        max_replication_slots: 10

  initdb:
    - encoding: UTF8
    - data-checksums

  pg_hba:
    - host replication replicator 10.0.0.0/24 md5
    - host all all 0.0.0.0/0 md5

postgresql:
  listen: 10.0.0.11:5432
  connect_address: 10.0.0.11:5432
  data_dir: /var/lib/postgresql/16/main
  bin_dir: /usr/lib/postgresql/16/bin
  pgpass: /var/lib/postgresql/.pgpass
  authentication:
    replication:
      username: replicator
      password: StrongReplicationPassword
    superuser:
      username: postgres
      password: StrongSuperuserPassword

After executing these configurations on all three nodes, initialize Patroni via a custom Systemd service unit. Once the primary node boots, it will register its state in Etcd, and the subsequent nodes will automatically clone the data using pg_basebackup to become active standbys.

Validation and Failover Testing

A high-availability architecture is only as good as its verified failover performance. To monitor your newly provisioned cluster, utilize the Patroni administration tool:

patronictl -c /etc/patroni/patroni.yml list

This command outputs a real-time matrix indicating the role, status, TL (Timeline), and replication lag of each node. To simulate an unexpected infrastructure crash, abruptly terminate the active primary node or stop its Patroni service:

sudo systemctl stop patroni

By observing the logs on the remaining nodes, you will witness the automated failover sequence in action:

  1. The standby nodes detect that the leader key lease in Etcd has expired.
  2. The standby nodes evaluate their current replication timelines and log sequence numbers (LSN).
  3. The most advanced standby node claims the leader key and promotes itself to primary.
  4. Client connections are redirected to the new primary node via a load balancer or virtual IP.

Conclusion and Operational Best Practices

Implementing Patroni and Etcd on Cloud Servers successfully elevates your PostgreSQL infrastructure to enterprise-level availability. However, building the cluster is only the initial phase. For long-term operational success, observe the following best practices:

  • Implement Connection Pooling: Integrate a tool like PgBouncer coupled with HAProxy or a Cloud Load Balancer to abstract the cluster topology from your application layer, ensuring seamless transitions during switchovers.
  • Enable Automated Backups: High Availability provides redundancy against infrastructure failure, not data corruption. Continue to perform regular logical and physical backups (e.g., using pgBackRest).
  • Monitor Key Metrics: Establish monitoring alerts for Etcd disk write latencies and Patroni loop execution times to preemptively identify performance bottlenecks.

By adhering to this architectural framework, your organization can confidently deploy business-critical applications backed by a resilient, self-healing database tier capable of withstanding unexpected cloud infrastructure disruptions.

Building a Resilient Infrastructure: Configuring a High-Availability PostgreSQL Cluster with Patroni and Etcd on Cloud Servers | DPTCloud