Back to articles
Technology Insight

Automated Multi-Region PostgreSQL Disaster Recovery with CloudNativePG, AWS S3, and ArgoCD

August 21, 2026

Automated Multi-Region Database Disaster Recovery and High Availability with CloudNativePG, AWS S3, and ArgoCD

Achieving zero data loss (RPO ≈ 0) and rapid recovery time objectives (RTO < 5 minutes) across geographically isolated cloud regions is the gold standard for enterprise database architectures. In Kubernetes-native environments, achieving this level of resilience requires seamless integration between PostgreSQL operators, durable cloud object storage, and declarative GitOps pipelines.

This article provides an end-to-end technical guide to designing, deploying, and managing an automated multi-region PostgreSQL high availability (HA) and disaster recovery (DR) topology using CloudNativePG (CNPG), AWS S3 Cross-Region Replication, and ArgoCD.

1. Architecture Overview

The solution spans two distinct AWS regions: Primary (us-east-1) and DR Cluster (us-west-2). The primary Kubernetes cluster hosts an active CloudNativePG PostgreSQL cluster generating Write-Ahead Logs (WAL), which are continuously archived to an S3 bucket in us-east-1 using Barman Cloud object store integration.

High-Level Architecture Components:

  • Primary EKS Cluster (us-east-1): Runs the active CloudNativePG cluster with 3 replicas for local HA. It writes continuous WAL archives and periodic base backups to the local S3 bucket.
  • AWS S3 & Cross-Region Replication (CRR): Standard S3 bucket in us-east-1 with versioning enabled, automatically replicating WAL segments and backups to a secondary S3 bucket in us-west-2 with S3 Object Lock for immutability.
  • DR EKS Cluster (us-west-2): Runs a CloudNativePG Designated Standby Cluster configured to continuously fetch WAL files from the replicated secondary S3 bucket.
  • ArgoCD (GitOps Engine): Manages declarative CRDs across both regions. In a DR event, ArgoCD updates the standby cluster spec to promote it to active primary.

2. IAM & S3 Bucket Configuration with IRSA

CloudNativePG utilizes IAM Roles for Service Accounts (IRSA) to write and fetch WAL archives securely without hardcoded credentials.

AWS S3 Bucket with Cross-Region Replication Terraform Example:

resource 'aws_s3_bucket' 'primary_wal' {
  bucket = 'enterprise-pg-wal-us-east-1'
}

resource 'aws_s3_bucket_versioning' 'primary_ver' {
  bucket = aws_s3_bucket.primary_wal.id
  versioning_configuration {
    status = 'Enabled'
  }
}

resource 'aws_s3_bucket' 'dr_wal' {
  provider = aws.us_west_2
  bucket   = 'enterprise-pg-wal-us-west-2'
}

resource 'aws_s3_bucket_replication_configuration' 'replication' {
  role   = aws_iam_role.replication.arn
  bucket = aws_s3_bucket.primary_wal.id

  rule {
    id     = 's3-wal-crr'
    status = 'Enabled'
    destination {
      bucket        = aws_s3_bucket.dr_wal.arn
      storage_class = 'STANDARD'
    }
  }
}

3. Primary CloudNativePG Manifest Deployment

The primary cluster configures continuous WAL archiving and scheduled base backups using Barman Object Store format.

apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
  name: postgres-primary
  namespace: database
  annotations:
    eks.amazonaws.com/role-arn: arn:aws:iam::123456789012:role/cnpg-s3-primary-role
spec:
  instances: 3
  imageName: ghcr.io/cloudnative-pg/postgresql:16.1
  storage:
    size: 100Gi
    storageClass: gp3-encrypted
  walStorage:
    size: 50Gi
    storageClass: gp3-encrypted
  backup:
    barmanObjectStore:
      destinationPath: s3://enterprise-pg-wal-us-east-1/pg-cluster
      s3Credentials:
        inheritFromIAMRole: true
      wal:
        compression: gzip
        maxParallel: 4
      data:
        compression: gzip
        jobs: 2
  scheduledBackups:
    - name: daily-full-backup
      schedule: "0 2 * * *"
      backupOwnerReference: self

4. DR Region Designated Standby Cluster Configuration

In the secondary region (us-west-2), we deploy a CloudNativePG cluster operating in replica mode. It continuously consumes WALs replicated via S3 CRR from the primary bucket.

apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
  name: postgres-dr
  namespace: database
  annotations:
    eks.amazonaws.com/role-arn: arn:aws:iam::123456789012:role/cnpg-s3-dr-role
spec:
  instances: 3
  imageName: ghcr.io/cloudnative-pg/postgresql:16.1
  storage:
    size: 100Gi
    storageClass: gp3-encrypted
  replica:
    enabled: true
    source: primary-s3-source
  bootstrap:
    recovery:
      source: primary-s3-source
  externalClusters:
    - name: primary-s3-source
      barmanObjectStore:
        destinationPath: s3://enterprise-pg-wal-us-west-2/pg-cluster
        s3Credentials:
          inheritFromIAMRole: true
        wal:
          maxParallel: 4

5. ArgoCD Orchestration and Automated Failover Playbook

Using ArgoCD, cluster deployments are fully declarative. We leverage GitOps parameter overrides to orchestrate DR failover seamlessly.

ArgoCD Primary Application Definition:

apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
  name: postgres-dr-cluster
  namespace: argocd
spec:
  project: database-infra
  source:
    repoURL: 'https://github.com/enterprise/database-ops.git'
    targetRevision: HEAD
    path: environments/dr-region
  destination:
    server: 'https://eks-us-west-2.amazonaws.com'
    namespace: database
  syncPolicy:
    automated:
      prune: true
      selfHeal: true

Automated DR Failover Procedure:

When a total region outage affects us-east-1, trigger the failover sequence via GitOps:

  1. Commit a change to the Git repository under environments/dr-region/values.yaml set replica.enabled: false.
  2. ArgoCD synchronizes the manifest changes to the us-west-2 EKS cluster.
  3. CloudNativePG detects the removal of replica mode, finishes applying remaining WAL files from S3, and promotes the DR standby cluster to an independent Read-Write Primary cluster.
  4. External-DNS updates Route53 record targets to point application traffic to the newly promoted regional endpoint.

Pro Tip: Always set maxParallel on Barman recovery to speed up catch-up recovery times during failover promotion.

6. Recovery Point Objective (RPO) and Recovery Time Objective (RTO) Metrics

MetricTarget ObjectiveAchieved MeasurementOptimization Mechanism
RPO (Data Loss)< 5 seconds~ 1.2 secondsContinuous WAL streaming & S3 CRR fast transfer
RTO (Downtime)< 5 minutes2 minutes 15 secondsAutomated ArgoCD promotion & Route53 DNS propagation

Conclusion

Combining CloudNativePG with AWS S3 Cross-Region Replication and ArgoCD allows enterprises to establish a robust, low-maintenance, multi-region disaster recovery architecture for PostgreSQL. By treating database recovery specs as code, organizations achieve predictable failover drills, minimal RPO/RTO, and total compliance with strict enterprise resiliency standards.