Back to articles
Technology Insight

Optimizing Large-Scale Database Backups with Baikal and Wal-G: A Strategy for Minute-by-Minute Incremental Backups

June 1, 2026

Introduction: The Enterprise Challenge of Large-Scale Database Backups

In the modern data-driven economy, uptime and data integrity are the twin pillars of enterprise infrastructure. As databases scale into tens of terabytes, traditional backup methodologies rapidly disintegrate. Standard full backups consume prohibitive amounts of network bandwidth, trigger massive I/O bottlenecks, and take hours—sometimes days—to complete. For high-transaction systems, a daily or even hourly backup window is no longer sufficient. Enterprise risk management demands a drastically reduced Recovery Point Objective (RPO) and Recovery Time Objective (RTO).

To achieve a near-zero RPO without degrading production database performance, infrastructure engineers are turning to a powerful combination: Baikal (a distributed, HTAP-capable database cluster system built on top of MySQL protocols) and Wal-G (the successor to Wal-E, a highly efficient archival and restoration tool for block-based and write-ahead log backups). This technical guide explores how to design, implement, and optimize a minute-by-minute incremental backup strategy using Baikal and Wal-G.

---

Understanding the Core Technologies: Baikal and Wal-G

Before diving into the integration strategy, it is essential to understand why these two technologies complement each other so effectively in high-throughput environments.

What is Baikal?

Baikal is an enterprise-grade distributed database system designed to handle massive volumes of relational data while maintaining MySQL compatibility. It splits data into multiple shards (regions) and distributes them across a cluster. Because it handles high-concurrency read/write workloads across distributed nodes, capturing a consistent state across the entire cluster requires a sophisticated coordination mechanism rather than a naive file-copy approach.

What is Wal-G?

Wal-G is an open-source backup tool optimized for speed and resource efficiency. Originally built for PostgreSQL but heavily extended for other database engines, Wal-G leverages multi-threaded compression (LZ4, ZSTD) and parallelized uploads to cloud storage (such as AWS S3, Google Cloud Storage, or MinIO). It operates on two fundamental principles:

  • Delta Backups (Full/Incremental Base Backups): Only copying modified blocks of data since the last backup.
  • Continuous Archiving (Write-Ahead Logs/Binlogs): Shipping transaction logs to storage the moment they are finalized, enabling Point-in-Time Recovery (PITR).
---

Architecting the Minute-by-Minute Incremental Backup Strategy

Achieving minute-by-minute incremental backups requires a dual-track architecture: continuous transaction log shipping combined with frequent, lightweight delta backups. Below is the blueprint for this optimization strategy.

1. The Dual-Track Backup Pipeline

To achieve a 1-minute RPO efficiently, the backup pipeline must be decoupled into two parallel operations:

  1. The Continuous Log Stream: As Baikal nodes execute transactions, the underlying binary logs (or write-ahead logs depending on the specific storage engine layer like RocksDB/MyRocks) are continuously captured and pushed to object storage by Wal-G every 60 seconds or upon log rotation.
  2. The Scheduled Delta Backup: A daily or weekly base backup is established, followed by highly compressed, multi-threaded incremental delta backups taken at strategic intervals. Wal-G compares the data blocks against the previous backup and uploads only the modified blocks.
Key Insight: By shipping transaction logs every minute, you guarantee that in the event of a catastrophic failure, your maximum data loss is limited to less than 60 seconds of transactions.

2. Storage Layer Optimization

Writing backups directly to disk on the database server creates extreme I/O contention. Wal-G bypasses this by streaming compressed data blocks directly to centralized object storage. To support minute-by-minute operations, the storage bucket must be configured with a strict lifecycle policy and optimized for high parallel write requests.

---

Step-by-Step Implementation and Configuration

Setting up Wal-G to handle minute-by-minute incremental updates for a Baikal-managed architecture involves configuring the storage back-end, setting up the WAL/Binlog archiving intervals, and defining the delta thresholds.

Step 1: Environmental Configuration

Wal-G relies on environment variables or configuration files to define its behavior. Below is an optimized production configuration layout for handling high-frequency backups securely:

  • WALG_STREAM_COMPRESSION_METHOD: Set to zstd or lz4. Zstd offers superior compression ratios for relational data, minimizing network bandwidth, while LZ4 provides the lowest CPU overhead.
  • WALG_COMPRESSION_LEVEL: Tuned to strike a balance between CPU consumption and storage savings (typically level 3 for Zstd).
  • WALG_DELTA_MAX_STEPS: Defines how many incremental backups can chain together before forcing a new full base backup. For a high-frequency setup, a value of 7 to 14 is standard.

Step 2: Configuring the Minute-by-Minute Log Shipper

To ensure logs are pushed every minute, the database log archiver or a dedicated system daemon (such as a systemd timer or cron job) must trigger Wal-G's log-push command. If the volume of transactions is massive, Wal-G will automatically break the logs into smaller chunks, uploading them in parallel streams.

---

Critical Performance Tuning for Large Databases

Deploying a minute-by-minute backup on a multi-terabyte database without proper tuning will lead to resource exhaustion. Implement the following optimizations to safeguard production performance:

Throttling Disk I/O and Network Bandwidth

Because backups occur continuously, you must prevent Wal-G from consuming 100% of the disk read throughput. Use the WALG_DISK_RATE_LIMIT parameter to restrict the maximum bytes per second Wal-G can read from the disk array. Similarly, utilize cgroups or network traffic shaping to ensure backup uploads do not starve client application traffic.

Parallelism Settings

Wal-G allows you to define the number of concurrent upload workers using WALG_UPLOAD_CONCURRENCY. For large Baikal nodes leveraging NVMe storage, increasing this value allows Wal-G to utilize multiple CPU cores to compress and upload chunks simultaneously, significantly reducing the duration of delta calculations.

---

Disaster Recovery: Executing Point-in-Time Recovery (PITR)

A backup strategy is only as good as its restoration process. In a disaster scenario, Wal-G's architecture enables rapid restoration through a structured three-step process:

  1. Restoring the Base Backup: Wal-G fetches the closest full base backup from object storage and extracts it to the target data directory.
  2. Applying Delta Backups: Wal-G sequentially applies the incremental delta changes up to the latest possible block-level backup before the target recovery time.
  3. Replaying Transaction Logs: Finally, Wal-G fetches the continuous logs shipped minute-by-minute, replaying every transaction up to the exact second specified by the operator.

Because Wal-G utilizes highly parallelized download and decompression algorithms, the RTO is drastically shorter compared to traditional logical restores (such as replaying SQL dumps), making it highly compatible with enterprise SLA requirements.

---

Conclusion and Best Practices

Optimizing large-scale database backups requires moving away from legacy, monolithic backup windows and embracing continuous, block-level data streaming. Combining the distributed resilience of Baikal with the speed and efficiency of Wal-G provides a modern framework capable of sustaining minute-by-minute incremental backups on production infrastructure.

As a final checklist, ensure your backup strategy includes automated restoration testing. A backup is only verified once it has been successfully restored on a separate staging environment. By implementing strict I/O throttling, choosing optimal compression algorithms, and maintaining a disciplined delta chain length, you can safeguard your enterprise data against catastrophic failure while preserving maximum production performance.

Optimizing Large-Scale Database Backups with Baikal and Wal-G: A Strategy for Minute-by-Minute Incremental Backups | DPTCloud