Automated Multi-Region PostgreSQL Disaster Recovery with CloudNativePG, AWS S3, and ArgoCD
Automated Multi-Region Database Disaster Recovery and High Availability
In enterprise cloud-native environments, relying on single-region Kubernetes clusters for stateful workloads presents a severe business continuity risk. Achieving true business continuity for relational databases requires a multi-region disaster recovery (DR) strategy with strict Recovery Time Objectives (RTO) and Recovery Point Objectives (RPO). This guide covers building a production-grade, multi-region PostgreSQL disaster recovery architecture using CloudNativePG (CNPG), AWS S3 Cross-Region Replication (CRR), and ArgoCD.
1. Multi-Region Active-Passive Architecture Overview
The solution uses an Active-Passive cross-region architecture across two isolated Amazon EKS clusters situated in separate AWS regions (e.g., us-east-1 as Primary and us-west-2 as Secondary/Standby).
- Primary Region (us-east-1): Hosts a read-write CloudNativePG cluster. Continuous Write-Ahead Log (WAL) archiving and physical base backups are continuously streamed to an active AWS S3 bucket using Barman Cloud Engine.
- Storage Replication Layer: AWS S3 Cross-Region Replication (CRR) asynchronously replicates WAL archives and base backups from the primary S3 bucket to a target S3 bucket in
us-west-2with S3 Object Lock and KMS Customer Managed Keys (CMK). - Secondary Region (us-west-2): Hosts a CloudNativePG Designated Standby Cluster. It continuously restores WAL files from the replicated local S3 bucket, staying in near-real-time synchronization with the primary.
- GitOps Deployment & Control: ArgoCD orchestrates state configuration using ApplicationSets and declarative sync hooks to manage failover mechanics and switch regional roles seamlessly.
2. Primary PostgreSQL Cluster Configuration
The primary CloudNativePG cluster is configured with native Barman Cloud integrations for automated physical backup and continuous WAL archiving. The manifest below specifies cluster topology, pod anti-affinity, and S3 credentials using IAM Roles for Service Accounts (IRSA).
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: pg-primary-cluster
namespace: database
spec:
instances: 3
primaryUpdateStrategy: Unsupervised
imageName: ghcr.io/cloudnative-pg/postgresql:16.2
storage:
size: 200Gi
storageClass: gp3-encrypted
walStorage:
size: 50Gi
storageClass: gp3-encrypted
resources:
requests:
cpu: "2"
memory: 8Gi
limits:
cpu: "4"
memory: 16Gi
postgresql:
parameters:
max_connections: "500"
shared_buffers: 2Gi
work_mem: 16MB
archive_timeout: "60s"
backup:
barmanObjectStore:
destinationPath: s3://company-pg-backups-primary-us-east-1/
endpointURL: https://s3.us-east-1.amazonaws.com
s3Credentials:
inheritFromIAMRole: true
wal:
compression: gzip
maxParallel: 8
data:
compression: gzip
jobs: 4
affinity:
podAntiAffinityRequirement: Preferred
topologyKey: topology.kubernetes.io/zone3. Standby Designated Cluster Setup
In the secondary EKS cluster (us-west-2), we deploy a CloudNativePG cluster marked with replica.enabled: true and a designated externalClusters configuration targeting the replicated S3 bucket.
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: pg-standby-cluster
namespace: database
spec:
instances: 3
imageName: ghcr.io/cloudnative-pg/postgresql:16.2
replica:
enabled: true
source: primary-s3-source
storage:
size: 200Gi
storageClass: gp3-encrypted
externalClusters:
- name: primary-s3-source
barmanObjectStore:
destinationPath: s3://company-pg-backups-replica-us-west-2/
endpointURL: https://s3.us-west-2.amazonaws.com
s3Credentials:
inheritFromIAMRole: true
wal:
maxParallel: 8
bootstrap:
recovery:
source: primary-s3-source
resources:
requests:
cpu: "2"
memory: 8Gi
limits:
cpu: "4"
memory: 16Gi4. ArgoCD GitOps Automation with ApplicationSet
To enforce consistency and avoid split-brain states during deployments, ArgoCD uses an ApplicationSet configured with a Matrix Generator. This binds target EKS cluster parameters to their designated region-specific configurations.
apiVersion: argoproj.io/v1alpha1
kind: ApplicationSet
metadata:
name: postgres-multiregion-appset
namespace: argocd
spec:
generators:
- matrix:
generators:
- clusters:
selector:
matchLabels:
tier: database-infrastructure
- list:
elements:
- region: us-east-1
clusterRole: primary
manifestPath: infrastructure/postgres/primary
- region: us-west-2
clusterRole: standby
manifestPath: infrastructure/postgres/standby
template:
metadata:
name: 'pg-{{clusterRole}}-{{name}}'
spec:
project: default
source:
repoURL: 'https://github.com/enterprise/database-ops.git'
targetRevision: HEAD
path: '{{manifestPath}}'
destination:
server: '{{server}}'
namespace: database
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true5. Automated Regional Failover & Promotion Protocol
In the event of a catastrophic failure in us-east-1, the execution of automated failover requires a precise sequence to prevent split-brain and minimize data loss:
- Isolate Primary Region: Revoke ingress routing to the primary database via AWS Route 53 Application Recovery Controller (ARC) routing controls.
- Promote Designated Standby: Update the Standby Cluster manifest in Git, setting
spec.replica.enabled: falseor removing thereplicablock entirely. - GitOps Sync Trigger: ArgoCD detects the Git commit and applies the updated manifest to the secondary cluster in
us-west-2. CloudNativePG terminates recovery mode, promotes the standby instance to primary, and opens read-write access. - DNS Traffic Switch: Update Route 53 DNS records or ARC control states to point application traffic to the newly promoted cluster service in
us-west-2.
Note: To test RTO/RPO limits without impacting production data, execute dry-run failover scenarios using Chaos Mesh to simulate cross-region latency or network partition before initiating automated failovers.
