Back to articles
Technology Insight

Scaling Infrastructure: Building an Automated 100-VPS Clone Cluster Using OpenTofu, Terragrunt, and GitOps

June 5, 2026

Introduction to Enterprise-Scale Infrastructure Automation

In modern cloud engineering, provisioning infrastructure at scale demands more than just basic automation scripts; it requires absolute predictability, strict state management, and rapid execution. When tasked with deploying a cluster of 100 identical Virtual Private Servers (VPS) for distributed testing, edge computing, or horizontal scaling, manual configuration becomes an operational liability. Relying on ad-hoc shell scripts or unmanaged Infrastructure as Code (IaC) inevitably leads to configuration drift, security vulnerabilities, and debugging nightmares.

To solve this challenge at an enterprise level, this guide explores a robust, production-ready framework combining OpenTofu (the open-source evolution of Terraform), Terragrunt (the premier tool for keeping IaC code DRY), and GitOps methodologies. By adopting this declarative stack, organizations can transition from fragmented deployments to a centralized, version-controlled architecture capable of cloning, modifying, and tearing down hundreds of server instances seamlessly.

The Core Architectural Pillars: OpenTofu, Terragrunt, and GitOps

Before diving into the implementation steps, it is essential to understand why this specific toolchain represents the gold standard for high-volume VPS deployments.

1. OpenTofu: The Open-Source Foundation

Following changes in licensing paradigms across the IaC landscape, OpenTofu has emerged as a reliable, fully open-source engine for resource provisioning. It operates identically to traditional declarative systems, utilizing the HashiCorp Configuration Language (HCL) to model infrastructure state. For a 100-VPS deployment, OpenTofu ensures that every network interface, storage volume, and compute instance is explicitly defined and recorded in a centralized state file.

2. Terragrunt: Orchestrating at Scale and Keeping Code DRY

While OpenTofu handles the what, Terragrunt manages the how. If you were to use pure OpenTofu to deploy 100 separate environments or highly segmented VPS clusters, you would quickly find yourself copying and pasting blocks of code, leading to maintenance bottlenecks. Terragrunt acts as a wrapper that allows engineers to write a infrastructure module once and inherit it across multiple environments or instances. It automatically manages backend configurations, minimizes code duplication (Keeping it DRY - Don't Repeat Yourself), and orchestrates dependency graphs between your modules.

3. GitOps: Treating Infrastructure as Application Code

GitOps shifts the operational control plane to a Git repository. Every infrastructure change—whether scaling up to 100 nodes, changing a firewall rule, or upgrading a CPU allocation—is initiated via a Git commit and validated through a Pull Request (PR). Once merged, an automated CI/CD pipeline triggers the Terragrunt workflow, ensuring that the live infrastructure perfectly mirrors the code residing in the production branch.

Architecting the Codebase Structure

A scalable deployment requires a highly organized repository structure. To decouple the infrastructure definitions from the specific environment configurations, we split our setup into two core components: Modules and Live Environments.

The Reusable Module Layer

The module defines the structural blueprint of a single VPS or a small cluster block. It configures variables for compute resources, operating system images, network mappings, and security groups. By keeping this layer highly parameterized, it remains completely agnostic of whether it is deploying one instance or one hundred.

The Terragrunt Live Hierarchy

The live directory structures the actual deployments using hierarchical configuration files (terragrunt.hcl). The directory tree looks similar to this:

  • root.hcl - Defines global backend states (e.g., S3/Consul buckets for locks) and global providers.
  • prod/ - Represents the production environment.
    • vps-cluster-1/terragrunt.hcl - Configures the first batch of automated clones.
    • vps-cluster-2/terragrunt.hcl - Configures subsequent batches with overridden parameters if needed.

By leveraging Terragrunt's include block, individual environment files remain incredibly compact, often containing fewer than twenty lines of code focusing strictly on variable inputs rather than structural definitions.

Step-by-Step Implementation for 100 VPS Autopilot Deployment

Let us break down the configuration required to execute this massive deployment safely and efficiently.

Step 1: Writing the Reusable OpenTofu Module

First, create an OpenTofu module that accepts variables for instance counting. Using the count or for_each meta-arguments within HCL allows a single resource block to scale horizontally based on an input variable.

Design Pattern Note: When scaling to 100 instances, utilize lists or maps of objects to dynamically assign distinct static IPs, hostnames, or geographic zones to each VPS clone, preventing networking conflicts.

Step 2: Configuring the Terragrunt Root Layout

In your root terragrunt.hcl, specify the remote state management system. This ensures that when executing concurrent modifications, state locking prevents data corruption. Terragrunt can automatically generate the backend S3 buckets or database locks on the fly, reducing manual bootstrap steps.

Step 3: Scaling via Matrix Generation

To provision 100 clones without writing 100 directory blocks, utilize Terragrunt's capability to read external data inputs or local JSON files. By maintaining a single inventory file detailing the specifications of your 100 targets, Terragrunt can loop through the collection, injecting parameters dynamically into the underlying OpenTofu module.

Enforcing GitOps Pipelines for Automation and Guardrails

With the code structured, the final piece of the puzzle is the automation pipeline. Engineers should never run terragrunt apply from their local machines when dealing with enterprise infrastructure.

  1. The Pull Request Workflow: An engineer modifies the inventory file to alter the VPS specs or add new clones. Upon submitting a PR, a GitHub Action or GitLab CI runner executes terragrunt plan. The output is posted directly into the PR comments, allowing the team to audit exactly what resources will be created, modified, or destroyed.
  2. Automated Compliance Verification: Integrate security linters and policy-as-code engines (such as Open Policy Agent or Checkov) into the pipeline. If any of the 100 VPS clones violate networking security protocols, the pipeline halts immediately.
  3. The Continuous Deployment Phase: Once approved and merged into the main branch, the pipeline executes terragrunt apply --terragrunt-non-interactive. The OpenTofu engine communicates concurrently with the target cloud provider or private hypervisor APIs, spinning up the 100 instances in parallel execution threads.

Operational Challenges and Mitigation Strategies

Deploying at this scale introduces unique engineering hurdles that must be anticipated:

  • API Rate Limiting: Sending requests to provision 100 virtual machines simultaneously can trigger rate limits on your cloud provider. To mitigate this, configure Terragrunt's concurrency limits using the --terragrunt-parallelism flag to batch allocations gracefully.
  • State File Bloat: Managing 100 complex instances in a single state file can slow down plan and apply phases. Consider splitting the 100 VPS clones into smaller logical groups (e.g., 5 blocks of 20 nodes) mapped to distinct Terragrunt directories to optimize execution speeds.
  • Post-Provisioning Configuration: OpenTofu excels at provisioning hardware but is not a configuration manager. Pair your GitOps pipeline with cloud-init or Ansible. Once OpenTofu brings up the VPS, cloud-init should take over to inject SSH keys, configure monitoring daemons, and register the nodes to your central network.

Conclusion

Building a self-contained OpenTofu and Terragrunt cluster managed via GitOps elevates infrastructure management from an unstable scripting exercise to a highly reliable, audited software engineering practice. Scaling to 100 automated VPS clones becomes as simple as modifying a configuration file and committing it to Git. By adopting this declarative, DRY, and pipeline-driven approach, your business gains unprecedented agility, bulletproof consistency, and an infrastructure capable of scaling effortlessly alongside operational demands.