Thảm Họa & Khôi Phục PostgreSQL Đa Vùng Tự Động Với CloudNativePG, AWS S3 và ArgoCD
Tự Động Hóa Khai Thác & Khôi Phục Thảm Họa Cơ Sở Dữ Liệu Đa Vùng
Trong môi trường Cloud-Native dành cho doanh nghiệp lớn, việc phụ thuộc vào một cụm Kubernetes đơn lẻ ở một vùng duy nhất để vận hành cơ sở dữ liệu quan hệ tiềm ẩn rủi ro rất cao. Để đảm bảo tính liên tục của ứng dụng (Business Continuity), việc thiết lập chiến lược Khôi phục thảm họa (Disaster Recovery - DR) đa vùng với chỉ số RTO (Recovery Time Objective) và RPO (Recovery Point Objective) gần như bằng 0 là ưu tiên hàng đầu. Bài viết này hướng dẫn chi tiết giải pháp triển khai hạ tầng khôi phục thảm họa PostgreSQL đa vùng sử dụng CloudNativePG (CNPG), AWS S3 Cross-Region Replication (CRR) và ArgoCD.
1. Kiến Trúc Tổng Quan Đa Vùng Active-Passive
Giải pháp áp dụng mô hình Active-Passive đa vùng triển khai trên 2 cụm Amazon EKS hoàn toàn độc lập đặt tại 2 vùng khác nhau của AWS (Ví dụ: us-east-1 là Vùng Chính / Primary và us-west-2 là Vùng Dự Phòng / Standby).
- Vùng Chính (us-east-1): Chạy cụm CloudNativePG cho phép Đọc/Ghi (Read-Write). Các tệp log ghi trước WAL (Write-Ahead Log) và bản sao lưu vật lý base-backup được đẩy liên tục vào S3 Bucket thông qua cơ chế Barman Cloud.
- Tầng Sao Lưu S3: Tính năng AWS S3 Cross-Region Replication (CRR) liên tục sao chép bất đồng bộ các WAL archive từ S3 bucket vùng chính sang S3 bucket phụ ở
us-west-2, kết hợp với S3 Object Lock và mã hóa bằng KMS Customer Managed Key (CMK). - Vùng Dự Phòng (us-west-2): Chạy cụm CloudNativePG ở chế độ Designated Standby Cluster. Cụm này liên tục đọc WAL archive từ S3 bucket bản sao ở vùng
us-west-2để khôi phục dữ liệu theo thời gian thực (near-real-time). - Quản Lý GitOps: ArgoCD đồng bộ hóa và quản lý cấu hình các cụm thông qua ApplicationSets và các hook chuyển đổi vai trò (Failover Switches) bằng mã nguồn khai báo.
2. Cấu Hình Cụm Primary PostgreSQL
Cụm CloudNativePG chính được tích hợp sẵn engine Barman Cloud để tự động lưu trữ WAL archive và thực hiện bản sao lưu định kỳ. Manifest bên dưới cấu hình topology cụm, pod anti-affinity và xác thực AWS S3 thông qua IAM Roles for Service Accounts (IRSA).
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: pg-primary-cluster
namespace: database
spec:
instances: 3
primaryUpdateStrategy: Unsupervised
imageName: ghcr.io/cloudnative-pg/postgresql:16.2
storage:
size: 200Gi
storageClass: gp3-encrypted
walStorage:
size: 50Gi
storageClass: gp3-encrypted
resources:
requests:
cpu: "2"
memory: 8Gi
limits:
cpu: "4"
memory: 16Gi
postgresql:
parameters:
max_connections: "500"
shared_buffers: 2Gi
work_mem: 16MB
archive_timeout: "60s"
backup:
barmanObjectStore:
destinationPath: s3://company-pg-backups-primary-us-east-1/
endpointURL: https://s3.us-east-1.amazonaws.com
s3Credentials:
inheritFromIAMRole: true
wal:
compression: gzip
maxParallel: 8
data:
compression: gzip
jobs: 4
affinity:
podAntiAffinityRequirement: Preferred
topologyKey: topology.kubernetes.io/zone3. Cấu Hình Cụm Standby (Designated Replica Cluster)
Tại cụm EKS phụ ở vùng us-west-2, chúng ta khai báo một cụm CloudNativePG được bật cờ replica.enabled: true và cấu hình externalClusters trỏ trực tiếp tới S3 Bucket đã được nhân bản ở vùng dự phòng.
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: pg-standby-cluster
namespace: database
spec:
instances: 3
imageName: ghcr.io/cloudnative-pg/postgresql:16.2
replica:
enabled: true
source: primary-s3-source
storage:
size: 200Gi
storageClass: gp3-encrypted
externalClusters:
- name: primary-s3-source
barmanObjectStore:
destinationPath: s3://company-pg-backups-replica-us-west-2/
endpointURL: https://s3.us-west-2.amazonaws.com
s3Credentials:
inheritFromIAMRole: true
wal:
maxParallel: 8
bootstrap:
recovery:
source: primary-s3-source
resources:
requests:
cpu: "2"
memory: 8Gi
limits:
cpu: "4"
memory: 16Gi4. Đồng Bộ GitOps Với ArgoCD ApplicationSet
Để đảm bảo tính nhất quán và ngăn ngừa hiện tượng xung đột dữ liệu (split-brain) khi triển khai, ArgoCD sử dụng ApplicationSet với Matrix Generator. Cấu hình này liên kết các tham số cụm Kubernetes mục tiêu với từng file khai báo riêng cho từng vùng.
apiVersion: argoproj.io/v1alpha1
kind: ApplicationSet
metadata:
name: postgres-multiregion-appset
namespace: argocd
spec:
generators:
- matrix:
generators:
- clusters:
selector:
matchLabels:
tier: database-infrastructure
- list:
elements:
- region: us-east-1
clusterRole: primary
manifestPath: infrastructure/postgres/primary
- region: us-west-2
clusterRole: standby
manifestPath: infrastructure/postgres/standby
template:
metadata:
name: 'pg-{{clusterRole}}-{{name}}'
spec:
project: default
source:
repoURL: 'https://github.com/enterprise/database-ops.git'
targetRevision: HEAD
path: '{{manifestPath}}'
destination:
server: '{{server}}'
namespace: database
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=true5. Quy Trình Chuyển Vùng Tự Động (Failover & Promotion)
Khi vùng chính us-east-1 gặp sự cố nghiêm trọng, quy trình chuyển vùng (Failover) cần thực hiện theo thứ tự chuẩn xác để đảm bảo an toàn dữ liệu:
- Cô lập vùng cũ: Ngắt luồng traffic truy cập vào cụm Primary thông qua AWS Route 53 Application Recovery Controller (ARC).
- Thăng cấp cụm Standby: Cập nhật file manifest của cụm Standby trên Git: đổi
spec.replica.enabled: falsehoặc loại bỏ phần cấu hìnhreplica. - Kích hoạt GitOps Sync: ArgoCD phát hiện thay đổi trên Git và áp dụng manifest mới xuống cụm EKS tại
us-west-2. CloudNativePG sẽ dừng tiến trình recovery, nâng cấp node Standby thành Primary và mở quyền Đọc/Ghi. - Chuyển hướng Traffic DNS: Cập nhật bản ghi DNS Route 53 để điều hướng toàn bộ kết nối của ứng dụng sang cụm PostgreSQL mới tại
us-west-2.
Lưu ý: Để kiểm thử chỉ số RTO/RPO định kỳ mà không ảnh hưởng tới dữ liệu sản xuất, doanh nghiệp nên kết hợp sử dụng Chaos Mesh để giả lập sự cố mất kết nối giữa các vùng trước khi thực thi chuyển vùng thực tế.
