Back to articles
Technology Insight

Building a Self-Healing GitOps Infrastructure: Combining Keepalived and K3s on Budget VPS Clusters

May 29, 2026

Introduction: The Challenge of High Availability on a Budget

In the modern DevOps landscape, GitOps has emerged as the gold standard for infrastructure automation and continuous delivery. By using a Git repository as the single source of truth, organizations can ensure consistency, traceability, and rapid recovery. However, implementing a robust GitOps workflow traditionally requires expensive cloud-managed Kubernetes services (like EKS or GKE) and managed load balancers to achieve high availability (HA).

For startups, independent developers, and small-to-medium enterprises (SMEs), these cloud costs can rapidly escalate. This guide explores a powerful, cost-effective alternative: building a self-healing GitOps infrastructure by combining Keepalived and K3s on a cluster of budget Virtual Private Servers (VPS). By leveraging these lightweight yet enterprise-grade tools, you can achieve automated failure recovery and infrastructure resilience at a fraction of the cost.

The Architecture: K3s, Keepalived, and GitOps Convergence

To understand how this system achieves self-healing capabilities, we must look at the synergy between its core components:

  • K3s by Rancher: A highly lightweight, fully compliant Kubernetes distribution designed for resource-constrained environments. It reduces the memory footprint of standard Kubernetes while retaining all essential APIs.
  • Keepalived: A routing software based on the Virtual Router Redundancy Protocol (VRRP). It provides a Floating Virtual IP (VIP) that automatically switches between master nodes if the primary node goes offline.
  • GitOps Controller (Argo CD or Flux): Running inside K3s, this controller constantly monitors the cluster state against the Git repository, automatically correcting any configuration drift.

When combined, Keepalived ensures network-level resilience (infrastructure layer HA), while K3s and the GitOps engine ensure application and configuration resilience (orchestration layer HA).

Step-by-Step Infrastructure Design

1. Setting Up Network Redundancy with Keepalived

The first point of failure in a budget VPS cluster is the control plane network interface. If the primary master node goes down, the entire GitOps pipeline halts. Keepalived solves this by assigning a single Virtual IP (VIP) to the cluster. All traffic—including GitOps Webhooks and API server requests—points to this VIP.

On each VPS node, Keepalived runs a background daemon. The configuration uses a priority system where the primary node holds the VIP. If the primary node fails to send VRRP advertisements, the backup node with the next highest priority instantly claims the VIP. The transition happens within seconds, ensuring zero downtime for external traffic.

2. Deploying a Highly Available K3s Cluster

With the network layer secured by Keepalived, we deploy K3s in an HA configuration. Instead of using an external database like PostgreSQL (which adds cost and complexity), we utilize K3s' embedded ETCD or Embedded HA via SQLite (kine).

Pro Tip: For a budget 3-node VPS setup, initializing K3s with the --cluster-init flag activates the embedded ETCD datastore, providing excellent fault tolerance without requiring external database provisioning.

When executing the K3s installation script, the --tls-san parameter must point to the Keepalived Virtual IP. This ensures that the Kubernetes API server accepts secure requests regardless of which physical VPS is currently hosting the control plane.

Implementing the Self-Healing GitOps Loop

Once the infrastructure is up, we install a GitOps controller like Argo CD. The self-healing magic happens through a continuous reconciliation loop consisting of three phases:

  1. Observation: The GitOps controller polls the Git repository and observes the live K3s cluster.
  2. Calculation: It calculates differences between the target state (Git) and the live state (K3s).
  3. Correction: If a discrepancy is found—such as a deleted deployment or modified service configuration—the controller automatically overwrites the live state to match Git.

If a VPS node crashes entirely, Keepalived shifts the VIP to a healthy node. K3s automatically reschedules orphaned pods onto remaining nodes, and Argo CD validates that the deployed applications perfectly align with the Git configuration. Human intervention is completely eliminated from the recovery process.

Cost and Performance Optimization Strategies

Running Kubernetes on cheap VPS nodes (e.g., 2GB RAM, 1-2 vCPUs) requires strict resource management. To maximize efficiency and maintain high performance, apply these best practices:

  • Disable Unused K3s Components: By default, K3s installs Traefik, Local Storage, and CoreDNS. If you prefer alternative tools, use the --disable traefik flag during installation to save critical RAM.
  • Configure Resource Limits: Always define limits and requests in your GitOps application manifests. This prevents a single malfunctioning application from causing a cascading failure across your budget nodes.
  • Optimize Keepalived Check Intervals: Set your VRRP check interval to 2-3 seconds. Setting it too low causes unnecessary CPU consumption on low-end VPS processors, while setting it too high delays failover.

Conclusion: Enterprise Resilience at Startup Cost

Achieving high availability and self-healing infrastructure no longer requires enterprise-level budgets. By combining the network-level agility of Keepalived with the lightweight power of K3s, and wrapping them in a strict GitOps paradigm, you can create a production-ready, fault-tolerant cluster on low-cost VPS instances.

This setup not only protects your applications from hardware failures but also ensures that your system configuration remains secure, auditable, and immutable through Git. For any organization looking to scale efficiently while controlling cloud spend, this architecture represents the ideal convergence of economy and resilience.

Building a Self-Healing GitOps Infrastructure: Combining Keepalived and K3s on Budget VPS Clusters | DPTCloud