Automated Multi-Region PostgreSQL Disaster Recovery with CloudNativePG, AWS S3, and ArgoCD
Automated Multi-Region PostgreSQL Disaster Recovery with CloudNativePG, AWS S3, and ArgoCD
Building high-availability PostgreSQL architectures across geographic regions requires reliable Write-Ahead Log (WAL) shipping, point-in-time recovery (PITR) capabilities, and declarative lifecycle management. CloudNativePG provides native Kubernetes abstractions for PostgreSQL, leveraging Barman Cloud for object storage synchronization. When combined with AWS S3 cross-region replication and ArgoCD, teams can implement an enterprise-grade GitOps DR pattern achieving Near-Zero Recovery Point Objective (RPO) and low Recovery Time Objective (RTO).
1. Multi-Region Architectural Topology
The architecture consists of a Primary Cluster located in us-east-1 (Region A) and a Designated Replica Cluster running in us-west-2 (Region B). Continuous WAL archives and base backups stream to an encrypted AWS S3 bucket in us-east-1, which cross-region replicates to an S3 bucket in us-west-2.
ArgoCD manages both Kubernetes clusters declaratively via GitOps, ensuring environment consistency and controlling promote/failover workflows through configuration changes.
2. AWS IAM and S3 Cross-Region Setup
Provision AWS IAM roles using IAM Roles for Service Accounts (IRSA) to grant CloudNativePG pods permission to upload and retrieve WALs without static credentials.
# IAM Trust Policy for CloudNativePG Operator ServiceAccount
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Principal": {
"Federated": "arn:aws:iam::123456789012:oidc-provider/oidc.eks.us-east-1.amazonaws.com/id/EXAMPLE123456"
},
"Action": "sts:AssumeRoleWithWebIdentity",
"Condition": {
"StringEquals": {
"oidc.eks.us-east-1.amazonaws.com/id/EXAMPLE123456:sub": "system:serviceaccount:postgres-system:cnpg-main"
}
}
}
]
}3. Primary Cluster Manifest (Region A: us-east-1)
The primary cluster writes WAL logs continuously to the primary S3 bucket using Barman Cloud Object Store integration natively supported by CloudNativePG.
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: postgres-primary
namespace: database
spec:
instances: 3
primaryUpdateStrategy: Unsupervised
storage:
size: 100Gi
storageClass: gp3-encrypted
walStorage:
size: 50Gi
storageClass: gp3-encrypted
backup:
barmanObjectStore:
destinationPath: s3://company-postgres-backups-us-east-1/primary-db
endpointURL: https://s3.us-east-1.amazonaws.com
s3Credentials:
inheritFromIAMRole: true
wal:
compression: gzip
maxParallel: 4
data:
compression: gzip
jobs: 2
resources:
requests:
cpu: "2"
memory: 4Gi
limits:
cpu: "4"
memory: 8Gi4. Designated Replica Cluster Manifest (Region B: us-west-2)
The replica cluster in us-west-2 runs in continuous standby mode, fetching WAL files from the local replicated S3 bucket. It maintains near-real-time synchronization without opening direct cross-region network links between database pods.
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: postgres-replica
namespace: database
spec:
instances: 3
replica:
enabled: true
source: primary-s3-source
externalClusters:
- name: primary-s3-source
barmanObjectStore:
destinationPath: s3://company-postgres-backups-us-west-2/primary-db
endpointURL: https://s3.us-west-2.amazonaws.com
s3Credentials:
inheritFromIAMRole: true
wal:
maxParallel: 4
storage:
size: 100Gi
storageClass: gp3-encrypted
resources:
requests:
cpu: "2"
memory: 4Gi
limits:
cpu: "4"
memory: 8Gi5. ArgoCD Application Management & Automated Failover
ArgoCD targets both primary and backup Kubernetes clusters using ApplicationSet resources. To execute an orchestrated disaster recovery failover:
- Step 1: Disable production traffic to
us-east-1via ExternalDNS / Route53 DNS switch. - Step 2: In Git, update the replica cluster spec to disable standby mode by removing
spec.replica.enabled: trueor setting it tofalse. - Step 3: Push the change to the repository. ArgoCD syncs the target state to
us-west-2, promoting the standby PostgreSQL cluster to a fully functional primary.
Operational Note: Ensure S3 Cross-Region Replication latency stays below 10 seconds to satisfy enterprise RPO requirements during emergency promotion scenarios.
