Quay lại danh sách
Tin tức công nghệ

Thảm Họa & Khôi Phục PostgreSQL Đa Vùng Tự Động Với CloudNativePG, AWS S3 và ArgoCD

21 tháng 8, 2026

Tự Động Hóa Khai Thác & Khôi Phục Thảm Họa Cơ Sở Dữ Liệu Đa Vùng

Trong môi trường Cloud-Native dành cho doanh nghiệp lớn, việc phụ thuộc vào một cụm Kubernetes đơn lẻ ở một vùng duy nhất để vận hành cơ sở dữ liệu quan hệ tiềm ẩn rủi ro rất cao. Để đảm bảo tính liên tục của ứng dụng (Business Continuity), việc thiết lập chiến lược Khôi phục thảm họa (Disaster Recovery - DR) đa vùng với chỉ số RTO (Recovery Time Objective) và RPO (Recovery Point Objective) gần như bằng 0 là ưu tiên hàng đầu. Bài viết này hướng dẫn chi tiết giải pháp triển khai hạ tầng khôi phục thảm họa PostgreSQL đa vùng sử dụng CloudNativePG (CNPG), AWS S3 Cross-Region Replication (CRR) và ArgoCD.

1. Kiến Trúc Tổng Quan Đa Vùng Active-Passive

Giải pháp áp dụng mô hình Active-Passive đa vùng triển khai trên 2 cụm Amazon EKS hoàn toàn độc lập đặt tại 2 vùng khác nhau của AWS (Ví dụ: us-east-1 là Vùng Chính / Primary và us-west-2 là Vùng Dự Phòng / Standby).

  • Vùng Chính (us-east-1): Chạy cụm CloudNativePG cho phép Đọc/Ghi (Read-Write). Các tệp log ghi trước WAL (Write-Ahead Log) và bản sao lưu vật lý base-backup được đẩy liên tục vào S3 Bucket thông qua cơ chế Barman Cloud.
  • Tầng Sao Lưu S3: Tính năng AWS S3 Cross-Region Replication (CRR) liên tục sao chép bất đồng bộ các WAL archive từ S3 bucket vùng chính sang S3 bucket phụ ở us-west-2, kết hợp với S3 Object Lock và mã hóa bằng KMS Customer Managed Key (CMK).
  • Vùng Dự Phòng (us-west-2): Chạy cụm CloudNativePG ở chế độ Designated Standby Cluster. Cụm này liên tục đọc WAL archive từ S3 bucket bản sao ở vùng us-west-2 để khôi phục dữ liệu theo thời gian thực (near-real-time).
  • Quản Lý GitOps: ArgoCD đồng bộ hóa và quản lý cấu hình các cụm thông qua ApplicationSets và các hook chuyển đổi vai trò (Failover Switches) bằng mã nguồn khai báo.

2. Cấu Hình Cụm Primary PostgreSQL

Cụm CloudNativePG chính được tích hợp sẵn engine Barman Cloud để tự động lưu trữ WAL archive và thực hiện bản sao lưu định kỳ. Manifest bên dưới cấu hình topology cụm, pod anti-affinity và xác thực AWS S3 thông qua IAM Roles for Service Accounts (IRSA).

apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
  name: pg-primary-cluster
  namespace: database
spec:
  instances: 3
  primaryUpdateStrategy: Unsupervised
  imageName: ghcr.io/cloudnative-pg/postgresql:16.2
  
  storage:
    size: 200Gi
    storageClass: gp3-encrypted

  walStorage:
    size: 50Gi
    storageClass: gp3-encrypted

  resources:
    requests:
      cpu: "2"
      memory: 8Gi
    limits:
      cpu: "4"
      memory: 16Gi

  postgresql:
    parameters:
      max_connections: "500"
      shared_buffers: 2Gi
      work_mem: 16MB
      archive_timeout: "60s"

  backup:
    barmanObjectStore:
      destinationPath: s3://company-pg-backups-primary-us-east-1/
      endpointURL: https://s3.us-east-1.amazonaws.com
      s3Credentials:
        inheritFromIAMRole: true
      wal:
        compression: gzip
        maxParallel: 8
      data:
        compression: gzip
        jobs: 4

  affinity:
    podAntiAffinityRequirement: Preferred
    topologyKey: topology.kubernetes.io/zone

3. Cấu Hình Cụm Standby (Designated Replica Cluster)

Tại cụm EKS phụ ở vùng us-west-2, chúng ta khai báo một cụm CloudNativePG được bật cờ replica.enabled: true và cấu hình externalClusters trỏ trực tiếp tới S3 Bucket đã được nhân bản ở vùng dự phòng.

apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
  name: pg-standby-cluster
  namespace: database
spec:
  instances: 3
  imageName: ghcr.io/cloudnative-pg/postgresql:16.2
  
  replica:
    enabled: true
    source: primary-s3-source

  storage:
    size: 200Gi
    storageClass: gp3-encrypted

  externalClusters:
    - name: primary-s3-source
      barmanObjectStore:
        destinationPath: s3://company-pg-backups-replica-us-west-2/
        endpointURL: https://s3.us-west-2.amazonaws.com
        s3Credentials:
          inheritFromIAMRole: true
        wal:
          maxParallel: 8

  bootstrap:
    recovery:
      source: primary-s3-source

  resources:
    requests:
      cpu: "2"
      memory: 8Gi
    limits:
      cpu: "4"
      memory: 16Gi

4. Đồng Bộ GitOps Với ArgoCD ApplicationSet

Để đảm bảo tính nhất quán và ngăn ngừa hiện tượng xung đột dữ liệu (split-brain) khi triển khai, ArgoCD sử dụng ApplicationSet với Matrix Generator. Cấu hình này liên kết các tham số cụm Kubernetes mục tiêu với từng file khai báo riêng cho từng vùng.

apiVersion: argoproj.io/v1alpha1
kind: ApplicationSet
metadata:
  name: postgres-multiregion-appset
  namespace: argocd
spec:
  generators:
    - matrix:
        generators:
          - clusters:
              selector:
                matchLabels:
                  tier: database-infrastructure
          - list:
              elements:
                - region: us-east-1
                  clusterRole: primary
                  manifestPath: infrastructure/postgres/primary
                - region: us-west-2
                  clusterRole: standby
                  manifestPath: infrastructure/postgres/standby
  template:
    metadata:
      name: 'pg-{{clusterRole}}-{{name}}'
    spec:
      project: default
      source:
        repoURL: 'https://github.com/enterprise/database-ops.git'
        targetRevision: HEAD
        path: '{{manifestPath}}'
      destination:
        server: '{{server}}'
        namespace: database
      syncPolicy:
        automated:
          prune: true
          selfHeal: true
        syncOptions:
          - CreateNamespace=true

5. Quy Trình Chuyển Vùng Tự Động (Failover & Promotion)

Khi vùng chính us-east-1 gặp sự cố nghiêm trọng, quy trình chuyển vùng (Failover) cần thực hiện theo thứ tự chuẩn xác để đảm bảo an toàn dữ liệu:

  1. Cô lập vùng cũ: Ngắt luồng traffic truy cập vào cụm Primary thông qua AWS Route 53 Application Recovery Controller (ARC).
  2. Thăng cấp cụm Standby: Cập nhật file manifest của cụm Standby trên Git: đổi spec.replica.enabled: false hoặc loại bỏ phần cấu hình replica.
  3. Kích hoạt GitOps Sync: ArgoCD phát hiện thay đổi trên Git và áp dụng manifest mới xuống cụm EKS tại us-west-2. CloudNativePG sẽ dừng tiến trình recovery, nâng cấp node Standby thành Primary và mở quyền Đọc/Ghi.
  4. Chuyển hướng Traffic DNS: Cập nhật bản ghi DNS Route 53 để điều hướng toàn bộ kết nối của ứng dụng sang cụm PostgreSQL mới tại us-west-2.

Lưu ý: Để kiểm thử chỉ số RTO/RPO định kỳ mà không ảnh hưởng tới dữ liệu sản xuất, doanh nghiệp nên kết hợp sử dụng Chaos Mesh để giả lập sự cố mất kết nối giữa các vùng trước khi thực thi chuyển vùng thực tế.