Self-Hosting Distributed Web Performance Testing with k6 and K3s: Simulating 50,000 Concurrent Users on a Budget VPS Cluster
Introduction: The Cost-Performance Dilemma in Load Testing
In modern software engineering, ensuring system reliability under peak traffic is a non-negotiable requirement. However, engineering teams frequently encounter a stark dilemma when attempting to execute high-scale load tests. Proprietary SaaS load testing platforms, while feature-rich, charge exorbitant premiums when scaling to tens of thousands of concurrent users. For startups and mid-market enterprises, running a single comprehensive load test simulating 50,000 concurrent users (CCUs) can cost thousands of dollars in platform credits.
Fortunately, the evolution of open-source cloud-native tooling offers a powerful alternative: self-hosting a distributed performance testing infrastructure. By combining k6 (an open-source, developer-centric load testing tool written in Go) with K3s (a highly optimized, lightweight Kubernetes distribution by Rancher), engineers can orchestrate massive load generation across a cluster of budget Virtual Private Servers (VPS). This guide provides a comprehensive, production-grade blueprint to architecting, deploying, and executing a distributed load test capable of simulating 50,000 CCUs for a fraction of the cost of commercial alternatives.
1. Architectural Blueprint: Why k6 and K3s?
Before diving into configuration, it is essential to understand why the combination of k6 and K3s represents the sweet spot for budget-conscious, high-performance infrastructure.
The k6 Efficiency Advantage
Traditional testing tools like Apache JMeter allocate a single thread per virtual user (VU). This multi-threading model introduces massive memory and CPU overhead due to context switching, often requiring substantial computing resources just to generate moderate load. Conversely, k6 utilizes an asynchronous, single-threaded JavaScript execution loop per OS thread via the Go runtime. This architecture enables a single k6 instance to efficiently manage thousands of VUs while consuming minimal memory, making it ideal for resource-constrained VPS environments.
K3s: Enterprise Orchestration on Lean Hardware
Standard Kubernetes (upstream K8s) is notoriously resource-heavy, often consuming a significant portion of a low-spec VPS's RAM just to maintain the control plane. K3s strips away legacy, alpha, and cloud-provider-specific plugins, packing the entire control plane into a single binary under 100 MB. This allows us to maximize the allocation of VPS hardware directly to load generation rather than cluster overhead.
Architecture Summary: We will deploy a K3s cluster consisting of one master node and multiple worker nodes across budget VPS instances. The k6 Operator will manage the distributed lifecycle, coordinating ak6-initializerjob and multiplek6-runnerpods to generate uniform, aggregated traffic against the target system.
2. Infrastructure Provisioning and Hardware Sizing
To simulate 50,000 concurrent users safely without bottlenecking our load generators, we must calculate resource requirements carefully. In k6, a standard VU script with basic HTTP requests typically consumes between 1 MB to 5 MB of memory, depending on script complexity, payload size, and response handling.
- Target CCU: 50,000
- Estimated Memory per VU: 2.5 MB (Average case with standard JSON responses)
- Total Required Cluster Memory: 50,000 * 2.5 MB = 125,000 MB (~122 GB)
- Safety Buffer (25%): ~150 GB of cluster-wide RAM
To optimize costs, we can distribute this load across budget VPS providers (such as Hetzner, DigitalOcean, or Linode). We will provision the following cluster topography:
- 1x Master Node (Control Plane): 4 vCPU, 8 GB RAM (To handle K3s API server, Grafana, and Prometheus telemetry).
- 5x Worker Nodes (Load Generators): 8 vCPU, 32 GB RAM each (Totaling 40 vCPUs and 160 GB RAM).
This layout ensures we have ample headroom to handle OS overhead and prevent network card throttling on individual instances.
3. Step-by-Step Cluster Deployment and Optimization
Step 3.1: Linux Kernel Tuning (Crucial for High Connections)
By default, Linux operating systems limit the number of open file descriptors and ephemeral ports, which will instantly choke a high-scale load test. Every concurrent connection requires a file descriptor. Run the following commands on all worker nodes to modify network stack variables:
sudo sysctl -w net.core.somaxconn=65535
sudo sysctl -w net.ipv4.ip_local_port_range="1024 65535"
sudo sysctl -w fs.file-max=2097152
echo "* soft nofile 1048576" | sudo tee -a /etc/security/limits.conf
echo "* hard nofile 1048576" | sudo tee -a /etc/security/limits.conf
sudo sysctl -p
Step 3.2: Installing K3s
On the Master Node, execute the K3s installation script while explicitly disabling unneeded components like Traefik to save resources:
curl -sfL [https://get.k3s.io](https://get.k3s.io) | sh -s - --disable traefik --disable local-storage
Retrieve the node token from the master node: cat /var/lib/rancher/k3s/server/node-token.
On each of the 5 Worker Nodes, join them to the cluster using the master node's IP and the retrieved token:
curl -sfL [https://get.k3s.io](https://get.k3s.io) | K3S_URL=https://:6443 K3S_TOKEN= sh -
4. Installing and Operating the k6 Operator
To coordinate the distributed test run automatically, we deploy the official Kubernetes k6 Operator. This controller monitors custom resources and handles the synchronization and partitioning of VUs across worker nodes.
Step 4.1: Deploy the Operator via Helm
helm repo add k6-charts [https://grafana.github.io/helm-charts](https://grafana.github.io/helm-charts)
helm repo update
helm install k6-operator k6-charts/k6-operator --namespace testing --create-namespace
Step 4.2: Writing the Distributed Test Manifest
We write our standard k6 test script inside a Kubernetes ConfigMap. Below is an abbreviated structure of the K6 Custom Resource Definition (CRD) that instructs the operator to spin up 25 parallel runner pods, with each pod executing 2,000 VUs to hit the cumulative 50,000 CCU target.
apiVersion: k6.io/v1alpha1
kind: K6
metadata:
name: distributed-perf-test
namespace: testing
spec:
parallelism: 25
script:
configMap:
name: k6-test-script
file: test.js
arguments: --tag test_id=load-50k-ccu
runner:
resources:
limits:
cpu: "1500m"
memory: "4Gi"
requests:
cpu: "1000m"
memory: "2Gi"
5. Centralized Telemetry and Performance Analysis
Running a distributed test yields disjointed logs unless structured metrics aggregation is in place. We configure k6 to stream metrics in real-time to a centralized Prometheus instance, visualized via a custom Grafana dashboard.
When executing the test, monitor three primary metrics vectors closely:
- http_req_duration (p95 and p99): Pinpoints the latency distribution. If the 99th percentile spikes dramatically while the average remains flat, your target system is experiencing micro-queuing or thread pool starvation.
- http_req_failed: The percentage of error responses (5xx errors indicate server crash; 504 indicates gateway timeouts; connection resets imply TCP queue overflows).
- Runner Resource Usage: Monitor your worker nodes via Prometheus
node_exporter. If load generator CPU utilization crosses 80%, the metrics collected may become skewed due to client-side bottlenecks.
Conclusion: High-Scale Assurance Without High-Scale Budgets
By leveraging open-source components—k6's highly performant Go runtime and K3s's minimal control plane footprint—engineering teams can successfully democratize performance testing. Building your own enterprise-grade distributed testing cluster on a low-cost VPS matrix breaks dependencies on restrictive SaaS pricing models, shifting the financial calculus of performance verification from a major line-item expense to a minor operational cost.
Implementing this infrastructure ensures your application is truly production-ready, giving your team the validation data required to handle high-concurrency traffic events with absolute confidence.
