Scaling Infrastructure with GitOps: Managing 50 VPS Instances Using OpenTofu, Terragrunt, and GitHub Actions
Introduction: The Challenge of Managing 50 VPS Instances
In modern cloud engineering, managing a handful of Virtual Private Servers (VPS) is relatively straightforward. However, as an organization scales to 50 or more VPS instances distributed across multiple cloud providers or region zones, manual configurations, custom bash scripts, and traditional configuration management tools rapidly become unmaintainable bottlenecks. The risks of configuration drift, security vulnerabilities, and deployment errors increase exponentially.
To solve this operational bottleneck, modern enterprise infrastructure has shifted toward GitOps—a paradigm where Git acts as the single source of truth for infrastructure state. By combining the open-source flexibility of OpenTofu, the multi-environment optimization of Terragrunt, and the continuous integration power of GitHub Actions, teams can establish an automated, scalable, and audit-ready infrastructure pipeline. This article details how to architect and implement this exact GitOps framework for 50 VPS instances.
The Core Architectural Pillars
Before diving into the code structure, it is critical to understand the specific role each tool plays within this GitOps ecosystem:
- OpenTofu: An open-source fork of Terraform, acting as the declarative Infrastructure as Code (IaC) engine. It defines the target state of your compute, network, and security resources.
- Terragrunt: A thin wrapper that keeps your OpenTofu code DRY (Don't Repeat Yourself). It allows you to define backend configurations and provider versions once, scaling identical configurations across 50 VPS instances seamlessly without duplicating code blocks.
- GitHub Actions: The automation engine that executes the GitOps pipeline. It triggers plans on Pull Requests and applies changes upon merging to the production branch, eliminating human error from the deployment process.
Designing a DRY Directory Structure
When dealing with 50 VPS instances, structuring your repository incorrectly will lead to maintenance nightmares. By leveraging Terragrunt, we can separate our immutable OpenTofu modules from our environment-specific variables. Below is the recommended enterprise-grade directory structure:
infrastructure-gitops/
├── .github/workflows/
│ └── gitops-pipeline.yml
├── modules/
│ └── vps-node/
│ ├── main.tf
│ ├── variables.tf
│ └── outputs.tf
└── live/
├── terragrunt.hcl (Root configuration)
├── production/
│ ├── region-us/
│ │ ├── app-server-1/
│ │ │ └── terragrunt.hcl
│ │ └── app-server-2/
│ │ └── terragrunt.hcl
│ └── env.hcl
└── staging/
└── env.hclIn this architecture, the modules/vps-node/ directory contains the standard OpenTofu manifests for provisioning a single VPS instance, along with its associated firewalls and storage blocks. The live/ directory contains only the Terragrunt configurations which inject unique parameters (such as IP blocks, instance sizes, and server names) into the shared module.
Implementing the Root Configuration
The root terragrunt.hcl file dynamically configures the remote state storage (e.g., an S3-compatible bucket or HashiCorp Consul backend) and provider locks for all sub-directories. This eliminates the need to manually copy-paste backend blocks across 50 separate directories.
Key Benefit: If you need to upgrade your OpenTofu provider version or change the backend bucket name, you only modify it once in the root file.
An example of the root terragrunt.hcl configuration configuration looks like this:
remote_state {
backend = "s3"
config = {
bucket = "my-company-gitops-state"
key = "${path_relative_to_include()}/tofu.tfstate"
region = "us-east-1"
encrypt = true
}
}The function path_relative_to_include() automatically tracks the directory structure, ensuring that live/production/region-us/app-server-1/ gets its own isolated state file in the bucket.
Automating with GitHub Actions
The final pillar of the GitOps workflow is automating execution. Infrastructure changes should never be executed from a local developer machine. Instead, engineers propose changes via Git Pull Requests, and GitHub Actions manages the lifecycle.
The GitOps Workflow Lifecycle
- The Proposal: An engineer modifies a
terragrunt.hclfile to upgrade a VPS specification and opens a Pull Request (PR) against the main branch. - The Validation: GitHub Actions triggers a workflow that executes
terragrunt run-all plan. The generated plan output is posted directly into the PR comments for peer review. - The Approval: A senior engineer reviews the plan and merges the PR.
- The Execution: GitHub Actions detects the merge on the main branch and runs
terragrunt run-all apply --terragrunt-non-interactiveto provision the changes across the affected VPS infrastructure.
Anatomy of the Workflow File
Below is a highly optimized GitHub Actions workflow configured to parse altered directories and execute Terragrunt operations efficiently using OpenTofu:
name: 'GitOps Infrastructure Pipeline'
on:
pull_request:
branches:
- main
push:
branches:
- main
jobs:
terragrunt:
name: 'Terragrunt Execution'
runs-on: ubuntu-latest
steps:
- name: Checkout Code
uses: actions/checkout@v4
- name: Setup OpenTofu
uses: opentofu/setup-opentofu@v1
with:
version: '1.8.0'
- name: Setup Terragrunt
uses: autero1/action-terragrunt@v3
with:
terragrunt-version: 'v0.55.0'
- name: Terragrunt Plan (PR Only)
if: github.event_name == 'pull_request'
run: terragrunt run-all plan --terragrunt-non-interactive
env:
AWS_ACCESS_KEY_ID: ${{ secrets.AWS_ACCESS_KEY_ID }}
AWS_SECRET_ACCESS_KEY: ${{ secrets.AWS_SECRET_ACCESS_KEY }}
- name: Terragrunt Apply (Merge Only)
if: github.event_name == 'push' && github.ref == 'refs/heads/main'
run: terragrunt run-all apply --terragrunt-non-interactive
env:
AWS_ACCESS_KEY_ID: ${{ secrets.AWS_ACCESS_KEY_ID }}
AWS_SECRET_ACCESS_KEY: ${{ secrets.AWS_SECRET_ACCESS_KEY }}Managing State Overlap and Scaling to 50 VPS
When running operations across 50 nodes, running a global configuration command can become slow. To optimize performance and maintain isolated blast radiuses, use Terragrunt's capability to isolate deployments. Instead of running a blanket run-all, the pipeline can be further optimized using tools like tj-actions/changed-files to target only the directories containing active modifications.
Additionally, state locking must be enabled (via DynamoDB or your backend's native lock manager) to prevent parallel GitHub Actions runs from corrupting the OpenTofu state files if multiple changes are merged simultaneously.
Conclusion: The Operational Dividend
Transitioning from manual individual VPS maintenance to a centralized GitOps workflow using OpenTofu, Terragrunt, and GitHub Actions delivers massive business value. It reduces provisioning times from hours to minutes, creates a historical audit trail via Git commit logs, and guarantees that your infrastructure configuration precisely matches production reality.
While the initial setup requires intentional planning around repository design and pipeline security, the operational payoff becomes clear when scaling. Whether you are managing 50 VPS instances or planning to grow to 500, this modular GitOps architecture ensures your operations remain secure, repeatable, and fundamentally scalable.
