Automated Multi-Region PostgreSQL Disaster Recovery with CloudNativePG, AWS S3, and ArgoCD
Automated Multi-Region Database Disaster Recovery and High Availability with CloudNativePG, AWS S3, and ArgoCD
Achieving zero data loss (RPO ≈ 0) and rapid recovery time objectives (RTO < 5 minutes) across geographically isolated cloud regions is the gold standard for enterprise database architectures. In Kubernetes-native environments, achieving this level of resilience requires seamless integration between PostgreSQL operators, durable cloud object storage, and declarative GitOps pipelines.
This article provides an end-to-end technical guide to designing, deploying, and managing an automated multi-region PostgreSQL high availability (HA) and disaster recovery (DR) topology using CloudNativePG (CNPG), AWS S3 Cross-Region Replication, and ArgoCD.
1. Architecture Overview
The solution spans two distinct AWS regions: Primary (us-east-1) and DR Cluster (us-west-2). The primary Kubernetes cluster hosts an active CloudNativePG PostgreSQL cluster generating Write-Ahead Logs (WAL), which are continuously archived to an S3 bucket in us-east-1 using Barman Cloud object store integration.
High-Level Architecture Components:
- Primary EKS Cluster (
us-east-1): Runs the active CloudNativePG cluster with 3 replicas for local HA. It writes continuous WAL archives and periodic base backups to the local S3 bucket. - AWS S3 & Cross-Region Replication (CRR): Standard S3 bucket in
us-east-1with versioning enabled, automatically replicating WAL segments and backups to a secondary S3 bucket inus-west-2with S3 Object Lock for immutability. - DR EKS Cluster (
us-west-2): Runs a CloudNativePG Designated Standby Cluster configured to continuously fetch WAL files from the replicated secondary S3 bucket. - ArgoCD (GitOps Engine): Manages declarative CRDs across both regions. In a DR event, ArgoCD updates the standby cluster spec to promote it to active primary.
2. IAM & S3 Bucket Configuration with IRSA
CloudNativePG utilizes IAM Roles for Service Accounts (IRSA) to write and fetch WAL archives securely without hardcoded credentials.
AWS S3 Bucket with Cross-Region Replication Terraform Example:
resource 'aws_s3_bucket' 'primary_wal' {
bucket = 'enterprise-pg-wal-us-east-1'
}
resource 'aws_s3_bucket_versioning' 'primary_ver' {
bucket = aws_s3_bucket.primary_wal.id
versioning_configuration {
status = 'Enabled'
}
}
resource 'aws_s3_bucket' 'dr_wal' {
provider = aws.us_west_2
bucket = 'enterprise-pg-wal-us-west-2'
}
resource 'aws_s3_bucket_replication_configuration' 'replication' {
role = aws_iam_role.replication.arn
bucket = aws_s3_bucket.primary_wal.id
rule {
id = 's3-wal-crr'
status = 'Enabled'
destination {
bucket = aws_s3_bucket.dr_wal.arn
storage_class = 'STANDARD'
}
}
}3. Primary CloudNativePG Manifest Deployment
The primary cluster configures continuous WAL archiving and scheduled base backups using Barman Object Store format.
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: postgres-primary
namespace: database
annotations:
eks.amazonaws.com/role-arn: arn:aws:iam::123456789012:role/cnpg-s3-primary-role
spec:
instances: 3
imageName: ghcr.io/cloudnative-pg/postgresql:16.1
storage:
size: 100Gi
storageClass: gp3-encrypted
walStorage:
size: 50Gi
storageClass: gp3-encrypted
backup:
barmanObjectStore:
destinationPath: s3://enterprise-pg-wal-us-east-1/pg-cluster
s3Credentials:
inheritFromIAMRole: true
wal:
compression: gzip
maxParallel: 4
data:
compression: gzip
jobs: 2
scheduledBackups:
- name: daily-full-backup
schedule: "0 2 * * *"
backupOwnerReference: self4. DR Region Designated Standby Cluster Configuration
In the secondary region (us-west-2), we deploy a CloudNativePG cluster operating in replica mode. It continuously consumes WALs replicated via S3 CRR from the primary bucket.
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: postgres-dr
namespace: database
annotations:
eks.amazonaws.com/role-arn: arn:aws:iam::123456789012:role/cnpg-s3-dr-role
spec:
instances: 3
imageName: ghcr.io/cloudnative-pg/postgresql:16.1
storage:
size: 100Gi
storageClass: gp3-encrypted
replica:
enabled: true
source: primary-s3-source
bootstrap:
recovery:
source: primary-s3-source
externalClusters:
- name: primary-s3-source
barmanObjectStore:
destinationPath: s3://enterprise-pg-wal-us-west-2/pg-cluster
s3Credentials:
inheritFromIAMRole: true
wal:
maxParallel: 45. ArgoCD Orchestration and Automated Failover Playbook
Using ArgoCD, cluster deployments are fully declarative. We leverage GitOps parameter overrides to orchestrate DR failover seamlessly.
ArgoCD Primary Application Definition:
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: postgres-dr-cluster
namespace: argocd
spec:
project: database-infra
source:
repoURL: 'https://github.com/enterprise/database-ops.git'
targetRevision: HEAD
path: environments/dr-region
destination:
server: 'https://eks-us-west-2.amazonaws.com'
namespace: database
syncPolicy:
automated:
prune: true
selfHeal: trueAutomated DR Failover Procedure:
When a total region outage affects us-east-1, trigger the failover sequence via GitOps:
- Commit a change to the Git repository under
environments/dr-region/values.yamlsetreplica.enabled: false. - ArgoCD synchronizes the manifest changes to the
us-west-2EKS cluster. - CloudNativePG detects the removal of replica mode, finishes applying remaining WAL files from S3, and promotes the DR standby cluster to an independent Read-Write Primary cluster.
- External-DNS updates Route53 record targets to point application traffic to the newly promoted regional endpoint.
Pro Tip: Always set
maxParallelon Barman recovery to speed up catch-up recovery times during failover promotion.
6. Recovery Point Objective (RPO) and Recovery Time Objective (RTO) Metrics
| Metric | Target Objective | Achieved Measurement | Optimization Mechanism |
|---|---|---|---|
| RPO (Data Loss) | < 5 seconds | ~ 1.2 seconds | Continuous WAL streaming & S3 CRR fast transfer |
| RTO (Downtime) | < 5 minutes | 2 minutes 15 seconds | Automated ArgoCD promotion & Route53 DNS propagation |
Conclusion
Combining CloudNativePG with AWS S3 Cross-Region Replication and ArgoCD allows enterprises to establish a robust, low-maintenance, multi-region disaster recovery architecture for PostgreSQL. By treating database recovery specs as code, organizations achieve predictable failover drills, minimal RPO/RTO, and total compliance with strict enterprise resiliency standards.
