Building a Resilient Multi-Primary PostgreSQL Architecture: Deploying Spilo and Patroni on Oracle Cloud ARM64 Free Tier
Introduction: The Quest for High Availability on a Budget
In the modern enterprise landscape, database uptime is non-negotiable. While cloud providers offer managed relational database services with high availability (HA) baked in, these solutions often come with prohibitive licensing and operational costs. For startups, independent developers, and enterprise innovation labs, finding a balance between robust fault tolerance and cost efficiency is a critical challenge.
Oracle Cloud Infrastructure (OCI) offers one of the most generous Free Tier programs in the industry, notably providing up to 4 Ampere Altra ARM64 vCPUs and 24 GB of RAM. By leveraging this compute resource alongside open-source powerhouse tools like Patroni and Spilo (Zalando's packaged PostgreSQL + Patroni Docker image), we can construct a production-grade, highly available PostgreSQL cluster. This guide walks through the architectural principles, configuration matrices, and deployment steps required to establish a resilient, multi-node PostgreSQL infrastructure across 3 VPS instances on OCI.
Understanding the Architecture: Patroni, Spilo, and Distributed Consensus
Before diving into configuration, it is essential to clarify a common terminology nuance. True 'Multi-Primary' (where every node can accept concurrent writes to the same dataset simultaneously without coordination) introduces immense complexity and replication conflict risks in traditional relational systems. Instead, what we are building is a Dynamic Single-Primary Cluster with Automated Failover. To the application layer, it behaves like a highly available multi-primary pool where write traffic is dynamically and seamlessly routed to the designated leader node.
The Role of Patroni
Patroni is an open-source template configured to manage PostgreSQL HA clusters using a Distributed Consensus Store (DCS) such as etcd or Consul. Patroni monitors the health of the local PostgreSQL instance, communicates with the DCS to maintain a 'leader' lock, and automatically reconfigures replicas when the primary node fails. This completely eliminates the dreaded split-brain scenario, where two nodes believe they are the primary simultaneously, causing catastrophic data corruption.
Why Spilo?
Developed and maintained by Zalando, Spilo bundles PostgreSQL, Patroni, and essential backup tools (like Wal-E/Wal-G) into a single, highly optimized Docker image. Using Spilo stream-lines deployment on ARM64 architectures, ensuring that all dependencies—including the specialized compilation requirements for Ampere Altra processors—are pre-packaged and thoroughly tested.
Infrastructure Prerequisites on Oracle Cloud Infrastructure
To establish a resilient cluster capable of handling a split-brain scenario gracefully, a minimum of three nodes is required to satisfy the Raft or Paxos consensus quorum ($2n+1$ fault tolerance rule). On OCI, you should provision the following:
- Compute Instances: 3x OCI Compute Shapes (Virtual Machines) running
Ubuntu 22.04 LTS (ARM64). Allocate 1 vCPU and 6 GB RAM to each node to evenly distribute the 4 vCPU / 24 GB Free Tier allotment (leaving 1 vCPU and 6 GB for an application server or load balancer). - Networking: A Virtual Cloud Network (VCN) with a subnet spanning across different Availability Domains (or Fault Domains if in a single-AD region).
- Security Lists / Firewalls: Ensure the following ports are open between the instances:
2379/tcpand2380/tcpfor etcd communication.8008/tcpfor Patroni REST API REST calls.5432/tcpfor native PostgreSQL database connections.
Step-by-Step Deployment Blueprint
Step 1: Setting Up the Distributed Consensus Store (etcd)
First, we must establish our source of truth. Install etcd on all three nodes to create a distributed key-value store cluster. On each node, run:
sudo apt-get update && sudo apt-get install -y etcdEdit the /etc/etcd/etcd.yml configuration file on each node. Ensure that the initial-cluster parameter contains the IP addresses of all three VPS nodes, and set the listen-peer-urls and listen-client-urls to look at your internal OCI private network interfaces. Once configured, restart the daemon:
sudo systemctl restart etcd && sudo systemctl enable etcdVerify cluster health by executing etcdctl endpoint health. All nodes must report a HEALTHY status before you proceed.
Step 2: Preparing Docker for ARM64
Since Spilo runs inside Docker, install the native ARM64 Docker Engine on all nodes:
sudo apt-get install -y docker.io
sudo systemctl enable --now dockerStep 3: Configuring and Launching Spilo Containers
Zalando provides multi-arch Docker images that support ARM64 out of the box. Create a deployment configuration shell script or use Docker Compose. Below is an example environment configuration block required for Spilo to initialize properly via Patroni:
docker run -d \
--name pg-cluster-node \
--net=host \
-e SCOPE=pg-cluster \
-e ETCD3_HOSTS=10.0.0.11:2379,10.0.0.12:2379,10.0.0.13:2379 \
-e POD_IP=10.0.0.11 \
-e PATRONI_POSTGRESQL_LISTEN=0.0.0.0:5432 \
-e PATRONI_POSTGRESQL_DATA=/home/postgres/pgdata/pgroot/data \
-e PATRONI_REPLICATION_USERNAME=replicator \
-e PATRONI_REPLICATION_PASSWORD=YourSecureReplicationPassword \
-e PATRONI_SUPERUSER_USERNAME=postgres \
-e PATRONI_SUPERUSER_PASSWORD=YourSecureSuperuserPassword \
-v /data/patroni:/home/postgres/pgdata \
zalando/spilo-15:3.0-p1Note: Ensure that you substitute POD_IP with the specific node's local OCI private IP address, and adjust the ETCD3_HOSTS list to reflect your exact topology.
Validation and Failover Testing
Once all three containers are active, Patroni will orchestrate a leader election. One node will safely initialize as the primary database, while the remaining two nodes will automatically clone the primary data state via PostgreSQL streaming replication and transition into read-only replicas.
Checking Cluster State
To inspect the real-time layout of your high-availability cluster, execute the Patroni management command line utility within any of your running Spilo containers:
docker exec -it pg-cluster-node patronictl -c /home/postgres/postgres.yml listThe resulting matrix should display one clear node designated as Leader with the status running, and two entries marked as Replica displaying active replication lag numbers near zero.
Simulating a Node Crash
To prove the resiliency of your new configuration, simulate an unexpected hardware crash on your active Leader node by hard-stopping its container or disabling its network interface:
docker stop pg-cluster-nodeMonitor the logs on the remaining healthy instances. Within seconds, the etcd cluster notices the loss of the leader lease heartbeat. Patroni instantly holds a new quorum election, selects the healthiest remaining replica with the most up-to-date Write-Ahead Log (WAL) position, promotes it to the new Leader, and reconfigured the final third node to pull replication streams from the new primary source.
Conclusion and Best Practices
By marrying Oracle Cloud's free-tier ARM64 computing instances with Zalando's enterprise-tested Spilo image, we successfully created a highly resilient, enterprise-grade database cluster for zero infrastructure cost. To ensure long-term stability of this deployment, always remember to maintain proper external volume backups (using WAL-G pointing to OCI Object Storage), implement a load balancer or connection proxy like PgBouncer coupled with HAProxy to route application connections smoothly during a failover event, and set up monitoring alerts for your etcd cluster health status.
