Building a Self-Healing GitOps Infrastructure: Combining Keepalived and K3s on Budget VPS Clusters
Introduction to High-Availability GitOps on a Budget
In the modern DevOps landscape, GitOps has emerged as the gold standard for continuous delivery and infrastructure management. By using Git repositories as the single source of truth, organizations can automate infrastructure provisioning and application deployment. However, setting up a highly available (HA) GitOps pipeline typically involves expensive cloud providers, managed Kubernetes services, and complex load balancers.
For small-to-medium businesses (SMBs), startups, and independent developers, these costs can quickly become prohibitive. Fortunately, enterprise-grade resilience does not have to come with an enterprise price tag. By combining Keepalived (a lightweight routing software) with K3s (Rancher’s highly optimized, lightweight Kubernetes distribution) on a cluster of budget Virtual Private Servers (VPS), you can build a robust, self-healing GitOps infrastructure that automatically recovers from node failures.
The Architecture: Keepalived meets K3s
To understand how this system achieves self-healing capabilities, we must look at the two core components powering the infrastructure layer:
- Keepalived (Virtual IP Failover): Keepalived uses the Virtual Router Redundancy Protocol (VRRP) to broadcast heartbeats between VPS nodes. It assigns a single Virtual IP (VIP) to the active master node. If the primary node goes offline, Keepalived instantly shifts the VIP to a healthy standby node, ensuring zero downtime for external traffic.
- K3s (Lightweight Kubernetes Cluster): K3s replaces heavy components of standard Kubernetes with lightweight alternatives, making it perfect for constrained VPS environments. By running K3s in an embedded database mode (using ETCD or external datastores like PostgreSQL), we create a multi-master control plane.
When integrated, Keepalived ensures that your GitOps controller (such as Argo CD or Flux CD) always has a stable endpoint to communicate with, even if the underlying physical server crashes. The GitOps agent continuously reconciles the state of the cluster with your Git repository, automatically correcting configuration drifts and restarting failed pods.
Step-by-Step Implementation Guide
Let us walk through the process of setting up this resilient architecture across three budget VPS nodes running Ubuntu Linux. For this guide, we assume the nodes have the following private IPs: 192.168.1.10, 192.168.1.11, and 192.168.1.12. The Shared Virtual IP (VIP) will be 192.168.1.100.
Step 1: Installing and Configuring Keepalived
First, install Keepalived on all three VPS instances to manage our floating Virtual IP:
sudo apt-get update && sudo apt-get install -y keepalived
Next, configure the primary node (/etc/keepalived/keepalived.conf) as the MASTER:
vrrp_instance VI_1 {
state MASTER
interface eth0
virtual_router_id 51
priority 101
advert_int 1
authentication {
auth_type PASS
auth_pass secret_password
}
virtual_ipaddress {
192.168.1.100
}
}
On the remaining two backup nodes, change the state to BACKUP and lower the priority value (e.g., 100 and 99). Start and enable the service across all nodes using sudo systemctl enable --now keepalived. You will see that the VIP (192.168.1.100) automatically binds to the master node.
Step 2: Deploying the Highly Available K3s Cluster
With the stable VIP in place, we can initialize our Kubernetes control plane. Run the following command on the primary node, ensuring we point the cluster API to our Virtual IP:
curl -sfL [https://get.k3s.io](https://get.k3s.io) | sh -s - server \
--cluster-init \
--tls-san 192.168.1.100
Once initialized, extract the node token located at /var/lib/rancher/k3s/server/node-token. Use this token and the VIP to join the other two nodes to the cluster as server masters:
curl -sfL [https://get.k3s.io](https://get.k3s.io) | sh -s - server \
--server [https://192.168.1.100:6443](https://192.168.1.100:6443) \
--token YOUR_EXTRACTED_TOKEN \
--tls-san 192.168.1.100
By registering all master nodes via the VIP, the Kubernetes control plane remains accessible even if the primary node goes completely dark.
Step 3: Setting Up the GitOps Automation Loop
With our highly available K3s cluster active, we can now install Argo CD to handle declarative GitOps automation. Deploy Argo CD using the official manifests:
kubectl create namespace argocd
kubectl apply -n argocd -f [https://raw.githubusercontent.com/argoproj/argo-cd/stable/manifests/install.yaml](https://raw.githubusercontent.com/argoproj/argo-cd/stable/manifests/install.yaml)
Point Argo CD to your Git repository containing your infrastructure manifests (Deployment files, Ingress paths, Network Policies). Enable Automated Pruning and Self-Healing options in your Argo CD Application configuration:
Pro Tip: Enabling automated self-healing ensures that if a user manually alters a Kubernetes resource or if an infrastructure glitch modifies configurations, Argo CD will instantly overwrite the changes to match the exact state defined in Git.
Testing the Self-Healing and Failover Mechanisms
An infrastructure architecture is only as good as its proven recovery capabilities. To test the self-healing and failover loops, we simulate a catastrophic hardware crash by abruptly stopping the primary VPS node hosting the Keepalived MASTER state.
- Network Layer Failover: Within milliseconds of the primary node going dark, Keepalived on the secondary backup node detects the missing VRRP heartbeat. It instantly claims the Virtual IP (
192.168.1.100). External APIs and developers accessing the cluster face no DNS propagation delays because the IP address itself remains unchanged. - Control Plane Recovery: Since K3s was configured with a distributed multi-master architecture, the remaining two nodes maintain quorum. The Kubernetes API continues to accept requests via the newly rerouted VIP.
- GitOps Re-synchronization: If any application pods were lost during the node crash, the K3s scheduler automatically redistributes them to the surviving nodes. Concurrently, Argo CD checks the state of the cluster against Git. If any microservices or configurations are missing, it triggers an immediate sync to restore operational stability.
Maximizing Efficiency on Low-Cost VPS
Running Kubernetes on budget VPS hosting (providers like Hetzner, Contabo, or DigitalOcean) requires careful resource optimization. Because budget instances often come with limited RAM (2GB to 4GB) and vCPUs, apply these optimization strategies:
- Disable Unused K3s Components: By default, K3s installs local storage provisions and a default Traefik ingress controller. If you prefer NGINX or external routing, disable them during installation using
--disable traefikto save valuable memory. - Resource Limits: Always define strict CPU and RAM limits within your GitOps Kubernetes deployment manifests. This prevents a single misbehaving application pod from crashing an entire low-cost VPS node.
- Tune Keepalived Check Intervals: Set your VRRP advertising intervals wisely. An advertisement interval of 1 second offers fast failovers without generating unnecessary internal network overhead on budget infrastructure links.
Conclusion
Building a resilient, high-availability platform does not require massive cloud budgets. By leveraging Keepalived for intelligent networking failover and K3s for resource-efficient container orchestration, you can establish an autonomous, self-healing GitOps loop on cheap VPS instances. This architecture guarantees that your deployments stay stable, your configurations remain consistent, and your applications stay online—giving you enterprise-grade reliability at a fraction of the cost.
