Back to articles
Technology Insight

Automated Multi-Region PostgreSQL Disaster Recovery with CloudNativePG, AWS S3, and ArgoCD

August 21, 2026

Automated Multi-Region Database Disaster Recovery and High Availability

In enterprise cloud-native environments, relying on single-region Kubernetes clusters for stateful workloads presents a severe business continuity risk. Achieving true business continuity for relational databases requires a multi-region disaster recovery (DR) strategy with strict Recovery Time Objectives (RTO) and Recovery Point Objectives (RPO). This guide covers building a production-grade, multi-region PostgreSQL disaster recovery architecture using CloudNativePG (CNPG), AWS S3 Cross-Region Replication (CRR), and ArgoCD.

1. Multi-Region Active-Passive Architecture Overview

The solution uses an Active-Passive cross-region architecture across two isolated Amazon EKS clusters situated in separate AWS regions (e.g., us-east-1 as Primary and us-west-2 as Secondary/Standby).

  • Primary Region (us-east-1): Hosts a read-write CloudNativePG cluster. Continuous Write-Ahead Log (WAL) archiving and physical base backups are continuously streamed to an active AWS S3 bucket using Barman Cloud Engine.
  • Storage Replication Layer: AWS S3 Cross-Region Replication (CRR) asynchronously replicates WAL archives and base backups from the primary S3 bucket to a target S3 bucket in us-west-2 with S3 Object Lock and KMS Customer Managed Keys (CMK).
  • Secondary Region (us-west-2): Hosts a CloudNativePG Designated Standby Cluster. It continuously restores WAL files from the replicated local S3 bucket, staying in near-real-time synchronization with the primary.
  • GitOps Deployment & Control: ArgoCD orchestrates state configuration using ApplicationSets and declarative sync hooks to manage failover mechanics and switch regional roles seamlessly.

2. Primary PostgreSQL Cluster Configuration

The primary CloudNativePG cluster is configured with native Barman Cloud integrations for automated physical backup and continuous WAL archiving. The manifest below specifies cluster topology, pod anti-affinity, and S3 credentials using IAM Roles for Service Accounts (IRSA).

apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
  name: pg-primary-cluster
  namespace: database
spec:
  instances: 3
  primaryUpdateStrategy: Unsupervised
  imageName: ghcr.io/cloudnative-pg/postgresql:16.2
  
  storage:
    size: 200Gi
    storageClass: gp3-encrypted

  walStorage:
    size: 50Gi
    storageClass: gp3-encrypted

  resources:
    requests:
      cpu: "2"
      memory: 8Gi
    limits:
      cpu: "4"
      memory: 16Gi

  postgresql:
    parameters:
      max_connections: "500"
      shared_buffers: 2Gi
      work_mem: 16MB
      archive_timeout: "60s"

  backup:
    barmanObjectStore:
      destinationPath: s3://company-pg-backups-primary-us-east-1/
      endpointURL: https://s3.us-east-1.amazonaws.com
      s3Credentials:
        inheritFromIAMRole: true
      wal:
        compression: gzip
        maxParallel: 8
      data:
        compression: gzip
        jobs: 4

  affinity:
    podAntiAffinityRequirement: Preferred
    topologyKey: topology.kubernetes.io/zone

3. Standby Designated Cluster Setup

In the secondary EKS cluster (us-west-2), we deploy a CloudNativePG cluster marked with replica.enabled: true and a designated externalClusters configuration targeting the replicated S3 bucket.

apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
  name: pg-standby-cluster
  namespace: database
spec:
  instances: 3
  imageName: ghcr.io/cloudnative-pg/postgresql:16.2
  
  replica:
    enabled: true
    source: primary-s3-source

  storage:
    size: 200Gi
    storageClass: gp3-encrypted

  externalClusters:
    - name: primary-s3-source
      barmanObjectStore:
        destinationPath: s3://company-pg-backups-replica-us-west-2/
        endpointURL: https://s3.us-west-2.amazonaws.com
        s3Credentials:
          inheritFromIAMRole: true
        wal:
          maxParallel: 8

  bootstrap:
    recovery:
      source: primary-s3-source

  resources:
    requests:
      cpu: "2"
      memory: 8Gi
    limits:
      cpu: "4"
      memory: 16Gi

4. ArgoCD GitOps Automation with ApplicationSet

To enforce consistency and avoid split-brain states during deployments, ArgoCD uses an ApplicationSet configured with a Matrix Generator. This binds target EKS cluster parameters to their designated region-specific configurations.

apiVersion: argoproj.io/v1alpha1
kind: ApplicationSet
metadata:
  name: postgres-multiregion-appset
  namespace: argocd
spec:
  generators:
    - matrix:
        generators:
          - clusters:
              selector:
                matchLabels:
                  tier: database-infrastructure
          - list:
              elements:
                - region: us-east-1
                  clusterRole: primary
                  manifestPath: infrastructure/postgres/primary
                - region: us-west-2
                  clusterRole: standby
                  manifestPath: infrastructure/postgres/standby
  template:
    metadata:
      name: 'pg-{{clusterRole}}-{{name}}'
    spec:
      project: default
      source:
        repoURL: 'https://github.com/enterprise/database-ops.git'
        targetRevision: HEAD
        path: '{{manifestPath}}'
      destination:
        server: '{{server}}'
        namespace: database
      syncPolicy:
        automated:
          prune: true
          selfHeal: true
        syncOptions:
          - CreateNamespace=true

5. Automated Regional Failover & Promotion Protocol

In the event of a catastrophic failure in us-east-1, the execution of automated failover requires a precise sequence to prevent split-brain and minimize data loss:

  1. Isolate Primary Region: Revoke ingress routing to the primary database via AWS Route 53 Application Recovery Controller (ARC) routing controls.
  2. Promote Designated Standby: Update the Standby Cluster manifest in Git, setting spec.replica.enabled: false or removing the replica block entirely.
  3. GitOps Sync Trigger: ArgoCD detects the Git commit and applies the updated manifest to the secondary cluster in us-west-2. CloudNativePG terminates recovery mode, promotes the standby instance to primary, and opens read-write access.
  4. DNS Traffic Switch: Update Route 53 DNS records or ARC control states to point application traffic to the newly promoted cluster service in us-west-2.

Note: To test RTO/RPO limits without impacting production data, execute dry-run failover scenarios using Chaos Mesh to simulate cross-region latency or network partition before initiating automated failovers.