Quay lại danh sách
Tin tức công nghệ

Tự động hóa DR và HA Database Đa Region với CloudNativePG, AWS S3 và ArgoCD

21 tháng 8, 2026

Kiến trúc Disaster Recovery & High Availability Database Đa Region Enterprise

Trong môi trường Kubernetes chạy các ứng dụng quan trọng (mission-critical), giải pháp High Availability (HA) trong một region đơn lẻ không còn đủ đáp ứng yêu cầu khắc nghiệt của doanh nghiệp. Một chiến lược Disaster Recovery (DR) vững chắc đòi hỏi sự kết hợp giữa các region địa lý độc lập. Bài viết này trình bày chi tiết cách thiết lập hệ thống PostgreSQL đa region tự động hoá khai báo sử dụng CloudNativePG, AWS S3 Cross-Region Replication (CRR) và ArgoCD.

Mô hình Kiến trúc System

Giải pháp kết nối hai cụm EKS độc lập nằm tại hai AWS region: us-east-1 (Primary Cluster) và us-west-2 (Designated Standby DR Cluster).

  • Primary Region (us-east-1): Vận hành cụm CloudNativePG 3 node active. Các bản ghi Write-Ahead Logs (WAL) và full backup được đẩy liên tục về S3 bucket local thông qua barman-cloud-wal-archive.
  • AWS S3 CRR: Tự động sao chép bất đồng bộ các file WAL binary từ bucket chính sang S3 bucket target tại us-west-2 kết hợp S3 Object Lock để chống tấn công ransomware.
  • DR Region (us-west-2): Vận hành cụm CloudNativePG ở chế độ replica (standby), liên tục đọc các đoạn WAL được đồng bộ sang bucket tại us-west-2.
  • ArgoCD GitOps: Quản lý toàn bộ trạng thái khai báo (declarative state) trên cả hai region, điều phối quá trình promote, failover và chuyển hướng lưu lượng truy cập chỉ qua các thao tác Git.

1. Thiết lập Hạ tầng AWS (S3 Replication & IRSA)

Cấu hình tính năng S3 Cross-Region Replication với IAM roles thông qua AWS CLI hoặc Terraform nhằm đảm bảo tốc độ đồng bộ WAL giữa hai bucket.

# Tạo Primary và Secondary S3 Bucket có bật Versioning
aws s3api create-bucket --bucket cnpg-wal-primary-us-east-1 --region us-east-1
aws s3api put-bucket-versioning --bucket cnpg-wal-primary-us-east-1 \
  --versioning-configuration Status=Enabled

aws s3api create-bucket --bucket cnpg-wal-dr-us-west-2 --region us-west-2 \
  --create-bucket-configuration LocationConstraint=us-west-2
aws s3api put-bucket-versioning --bucket cnpg-wal-dr-us-west-2 \
  --versioning-configuration Status=Enabled

# Cấu hình S3 Cross-Region Replication (CRR) Policy
cat <<EOF > replication-rules.json
{
  "Role": "arn:aws:iam::123456789012:role/s3-crr-replication-role",
  "Rules": [
    {
      "Status": "Enabled",
      "Priority": 1,
      "DeleteMarkerReplication": { "Status": "Disabled" },
      "Filter": {},
      "Destination": {
        "Bucket": "arn:aws:s3:::cnpg-wal-dr-us-west-2",
        "ReplicationTime": {
          "Status": "Enabled",
          "Time": { "Minutes": 15 }
        },
        "Metrics": {
          "Status": "Enabled",
          "EventThreshold": { "Minutes": 15 }
        }
      }
    }
  ]
}
EOF

aws s3api put-bucket-replication --bucket cnpg-wal-primary-us-east-1 \
  --replication-configuration file://replication-rules.json

2. Cấu hình Manifest CloudNativePG Primary

Triển khai cụm PostgreSQL chính tại us-east-1 tích hợp Barman Object Store thông qua IAM Roles for Service Accounts (IRSA).

apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
  name: postgres-primary
  namespace: database
spec:
  instances: 3
  primaryUpdateStrategy: Unsupervised
  
  storage:
    size: 100Gi
    storageClass: gp3-encrypted
    
  walStorage:
    size: 50Gi
    storageClass: gp3-encrypted

  resources:
    requests:
      cpu: "2"
      memory: "8Gi"
    limits:
      cpu: "4"
      memory: "16Gi"

  postgresql:
    parameters:
      shared_buffers: "2GB"
      work_mem: "32MB"
      max_connections: "500"
      archive_timeout: "60s"

  backup:
    barmanObjectStore:
      destinationPath: s3://cnpg-wal-primary-us-east-1/
      endpointURL: https://s3.us-east-1.amazonaws.com
      s3Credentials:
        inheritFromIAMRole: true
      wal:
        compression: gzip
        maxParallel: 8
      data:
        compression: gzip
        immediateCheckpoint: true
        jobs: 4
    retentionPolicy: "30d"

3. Cấu hình Cluster Designated Standby (Region DR)

Tại region us-west-2, triển khai cụm CloudNativePG được cấu hình dưới dạng replica node đọc trực tiếp từ S3 bucket đã được đồng bộ.

apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
  name: postgres-dr
  namespace: database
spec:
  instances: 3
  
  # Cấu hình chế độ Standby Replica
  replica:
    enabled: true
    source: s3-replicated-backup

  externalClusters:
    - name: s3-replicated-backup
      barmanObjectStore:
        destinationPath: s3://cnpg-wal-dr-us-west-2/
        endpointURL: https://s3.us-west-2.amazonaws.com
        s3Credentials:
          inheritFromIAMRole: true
        wal:
          maxParallel: 8

  storage:
    size: 100Gi
    storageClass: gp3-encrypted

  resources:
    requests:
      cpu: "2"
      memory: "8Gi"
    limits:
      cpu: "4"
      memory: "16Gi"

4. Tự động hóa GitOps với ArgoCD ApplicationSet

Sử dụng ArgoCD ApplicationSet kết hợp Matrix Generator để đồng bộ và kiểm soát manifest cho cả hai region một cách nhất quán.

apiVersion: argoproj.io/v1alpha1
kind: ApplicationSet
metadata:
  name: cloudnativepg-multiregion
  namespace: argocd
spec:
  generators:
    - list:
        elements:
          - cluster: primary-cluster-east
            region: us-east-1
            path: environments/us-east-1
          - cluster: dr-cluster-west
            region: us-west-2
            path: environments/us-west-2
  template:
    metadata:
      name: 'pg-{{region}}'
    spec:
      project: default
      source:
        repoURL: 'https://github.com/enterprise/database-infra.git'
        targetRevision: HEAD
        path: '{{path}}'
      destination:
        name: '{{cluster}}'
        namespace: database
      syncPolicy:
        automated:
          prune: true
          selfHeal: true
        syncOptions:
          - CreateNamespace=true

5. Quy trình Chuyển vùng Sự cố (DR Failover Playbook)

Trong trường hợp region us-east-1 gặp sự cố toàn diện, thực hiện quy trình nâng cấp (promotion) theo chuẩn GitOps:

Bước 1: Cô lập Region Primary (Ngăn chặn Split-Brain)
Cấu hình Fencing trên cụm Primary để chặn toàn bộ dữ liệu ghi phát sinh.

# Annotate primary cluster để kích hoạt tính năng Fencing
kubectl annotate cluster postgres-primary -n database \
  cnpg.io/fencedInstances='["*"]' --overwrite

Bước 2: Nâng cấp Standby Cluster trên GitOps
Cập nhật file environments/us-west-2/postgres.yaml trong Git repository để tắt chế độ replica:

# Cập nhật environments/us-west-2/postgres.yaml
spec:
  replica:
    enabled: false # Tắt chế độ standby và chuyển cluster thành Primary

Bước 3: Thực thi ArgoCD Sync & Kiểm tra Trạng thái

# Đồng bộ thay đổi ngay lập tức qua ArgoCD CLI
argocd app sync pg-us-west-2

# Kiểm tra trạng thái Cluster bằng CloudNativePG plugin
kubectl cnpg status postgres-dr -n database

So sánh Chỉ số Doanh nghiệp (Enterprise Metrics)

Chỉ số (Metric)HA Đơn Region Thông ThườngCloudNativePG Đa Region + CRR
RPO (Recovery Point Objective)~0 giây (Synchronous)< 60 giây (Asynchronous WAL S3 Sync)
RTO (Recovery Time Objective)< 30 giây< 3 phút (GitOps Promotion Sync)
Rủi ro Split-BrainThấp (Quorum nội bộ)Được kiểm soát nhờ Fencing & S3 Object Lock
Rủi ro Mất dữ liệu khi sập RegionRất cao (Mất cả Data Center)Tối thiểu (Nhờ S3 Replication Đa Region)

Kết luận

Sự kết hợp giữa CloudNativePG, AWS S3 Cross-Region Replication và ArgoCD mang lại cho doanh nghiệp giải pháp Disaster Recovery đa region toàn diện, sẵn sàng cho sản xuất và tuân thủ chặt chẽ nguyên tắc GitOps. Mô hình này giúp rút ngắn chỉ số RPO xuống dưới 60 giây đồng thời giảm thiểu sai sót thao tác thủ công khi có sự cố nghiêm trọng.