Building a Resilient K3s Control Plane: High Availability via Kube-VIP on Multi-VPS Architecture
Introduction: The Challenge of High Availability in Edge and Multi-VPS Environments
In the modern cloud-native landscape, Kubernetes has become the de facto standard for container orchestration. For small to medium enterprises, edge computing scenarios, and cost-conscious development teams, K3s—Rancher's lightweight Kubernetes distribution—offers a highly efficient, low-overhead alternative to upstream Kubernetes. However, transitioning a K3s cluster from a single-node development environment to a production-grade deployment demands High Availability (HA).
Traditionally, achieving HA for the Kubernetes Control Plane requires an external load balancer (such as an AWS ALB, HAProxy, or Nginx running on a dedicated instance) to distribute traffic across multiple master nodes. In a multi-VPS or bare-metal environment, this requirement introduces additional infrastructure costs, complexity, and another potential single point of failure. This is where Kube-VIP steps in. By leveraging Kube-VIP, you can construct a self-contained, highly available K3s control plane directly on your existing virtual private servers without relying on any external hardware or cloud-provider load balancing services.
Understanding the Core Architecture: K3s and Kube-VIP
Before diving into the implementation steps, it is crucial to understand how K3s and Kube-VIP interact to provide seamless failover capabilities.
The Role of Kube-VIP
Kube-VIP is an open-source, lightweight solution designed to provide a Virtual IP (VIP) and load balancing capabilities for Kubernetes clusters. It operates directly within the cluster network, utilizing either ARP (Layer 2) or BGP (Layer 3) protocols to broadcast the availability of the Virtual IP. In a standard multi-VPS setup where BGP configuration might not be accessible, Layer 2 mode is the preferred choice. Kube-VIP monitors the health of the control plane nodes; if the primary node hosting the VIP fails, Kube-VIP rapidly migrates the IP address to a healthy master node via leader election, ensuring uninterrupted access to the Kubernetes API server.
The Storage Layer: Embedded K3s etcd
To support a true HA control plane, the cluster state must be synchronized across all master nodes. K3s simplifies this requirement by providing an embedded etcd cluster. When initializing K3s in HA mode, etcd automatically replicates data across the master nodes, requiring a minimum of three nodes to achieve quorum and tolerate the loss of a single node.
Prerequisites and Environment Setup
To follow this guide successfully, ensure your environment meets the following baseline requirements:
- Three VPS Instances: Operating on the same private Local Area Network (LAN) with static IP addresses. For the purposes of this guide, we will use Ubuntu 22.04 LTS.
- IP Address Allocation: Each node needs a dedicated IP, and an unassigned IP on the same subnet must be reserved for the Virtual IP (VIP).
- Network Accessibility: Ensure that internal firewalls (e.g., UFW) permit traffic on required ports, specifically
6443(Kubernetes API),2379-2380(etcd), and7946(Kube-VIP keepalived/VRRP).
| Node Name | Role | IP Address |
|---|---|---|
| master-01 | Control Plane / etcd | 192.168.10.11 |
| master-02 | Control Plane / etcd | 192.168.10.12 |
| master-03 | Control Plane / etcd | 192.168.10.13 |
| Kube-VIP | Virtual IP (Cluster Endpoint) | 192.168.10.100 |
Step-by-Step Implementation Guide
Step 1: Generating the Kube-VIP Manifest
Because Kube-VIP must be active before the K3s API server starts on subsequent nodes, we will deploy it as a Static Pod on the first master node. First, connect to master-01 via SSH. We will pull the official Kube-VIP container image and generate the deployment manifest using the environment variables defined below:
export VIP=192.168.10.100
export INTERFACE=eth0
export KVVERSION=$(curl -s [https://api.github.com/repos/kube-vip/kube-vip/releases/latest](https://api.github.com/repos/kube-vip/kube-vip/releases/latest) | grep -Po '"tag_name": "\K[^"]*')
sudo mkdir -p /var/lib/rancher/k3s/agent/pod-manifests/
sudo docker run --rm ghcr.io/kube-vip/kube-vip:$KVVERSION \
manifest pod \
--interface $INTERFACE \
--address $VIP \
--controlplane \
--services \
--arp \
--leaderElection | sudo tee /var/lib/rancher/k3s/agent/pod-manifests/kube-vip.yamlNote: Ensure thatINTERFACEmatches the actual network interface name of your VPS (e.g.,eth0orens3), which can be verified using theip acommand.
Step 2: Initializing the First K3s Master Node
With the Kube-VIP static pod manifest securely positioned in the K3s manifests directory, we can initialize the first control plane node. The K3s installer will automatically detect the static pod and spin up Kube-VIP alongside the cluster components. Execute the following command on master-01:
curl -sfL [https://get.k3s.io](https://get.k3s.io) | sh -s - server \
--cluster-init \
--tls-san $VIP \
--write-kubeconfig-mode 644The --cluster-init flag instructs K3s to initialize the embedded etcd database, while the --tls-san $VIP flag ensures that the Kubernetes API server includes our Virtual IP in its TLS certificate validation pool. Once the installation script finishes, confirm that the Virtual IP is active by pinging 192.168.10.100 from your local machine or another server.
Step 3: Extracting the Cluster Token
To join the remaining master nodes to our newly created HA cluster, we need the security token generated by the initial node. Retrieve this token from master-01 using the following command:
sudo cat /var/lib/rancher/k3s/server/node-tokenCopy the output string to your clipboard; it will be required in the subsequent steps.
Step 4: Joining the Secondary Master Nodes
Now, shift your focus to master-02 and master-03. We will configure these nodes to join the cluster as control plane instances. Crucially, we point the installation command directly to our Virtual IP ($VIP) instead of individual node IPs. This guarantees that if master-01 experiences downtime, the installation and subsequent cluster management operations remain entirely unhindered.
Execute this command on both master-02 and master-03, ensuring you replace with the token retrieved in Step 3:
export VIP=192.168.10.100
export TOKEN=""
curl -sfL [https://get.k3s.io](https://get.k3s.io) | sh -s - server \
--server https://$VIP:6443 \
--token $TOKEN \
--tls-san $VIP \
--write-kubeconfig-mode 644 As these nodes join, Kube-VIP manifests must also be replicated onto them to allow them to participate in the leader election process. Copy the kube-vip.yaml file created in Step 1 from master-01 to the corresponding directory (/var/lib/rancher/k3s/agent/pod-manifests/) on both secondary nodes. This ensures that if the primary node defaults, the remaining nodes are fully equipped to assume control of the VIP instantly.
Verifying Cluster Health and Failover Mechanisms
With all three nodes initialized, it is time to verify the resilience of your new HA cluster. From any master node, run the following command to check the status of your nodes:
kubectl get nodes -o wideYou should see all three master nodes listed with a status of Ready, each indicating membership in the control plane and etcd roles.Executing a Chaos Test
To truly validate the High Availability mechanism, simulate a catastrophic node failure. You can do this by abruptly shutting down or rebooting master-01 (the current holder of the VIP). While executing the shutdown, continuously run a ping command against the VIP from your local machine:
ping 192.168.10.100You will observe a minimal drop in packet transmission—typically only 1 to 2 packets—before Kube-VIP detects the node's absence, orchestrates a new leader election among the remaining nodes, and remaps the VIP. Run kubectl get nodes via the VIP endpoint to confirm that the cluster remains fully responsive and manageable despite losing a core component.
Conclusion and Best Practices
By implementing Kube-VIP within a multi-VPS K3s architecture, you effectively unlock enterprise-grade High Availability without the financial overhead or operational complexity of external infrastructure. This self-contained approach is highly scalable, incredibly cost-efficient, and perfectly suited for cloud-agnostic architectures.
As you transition this setup into daily production use, remember to enforce standard operations management practices: implement regular etcd snapshot backups, monitor resource utilization across your master nodes, and ensure your private networking stack remains isolated and highly secure. With this setup, your applications can run confidently on a robust, failure-tolerant infrastructure foundation.
