Scaling to 50,000 Concurrent Users: Self-Hosting Distributed Web Performance Testing on a Budget with k6 and Kubernetes
Introduction: The Cost and Scale Dilemma of Performance Testing
In the modern digital economy, application performance is directly tied to business revenue. A sub-second delay in page load time can result in catastrophic drops in conversion rates, while a complete system outage during peak traffic events can inflict permanent brand damage. To mitigate these risks, rigorous performance testing is non-negotiable. However, organizations frequently encounter a significant roadblock: the prohibitive cost of enterprise load testing platforms.
Simulating high-volume traffic—such as 50,000 concurrent users (CUs)—typically demands massive infrastructure or expensive SaaS subscriptions that charge premium rates based on virtual user minutes. For startups, mid-sized enterprises, and budget-conscious engineering teams, these costs can render comprehensive testing unfeasible. Fortunately, open-source technology offers an elegant alternative. By combining k6, a developer-centric load testing tool, with Kubernetes (K8s) for container orchestration, businesses can self-host a highly scalable, distributed testing infrastructure on a cluster of affordable Virtual Private Servers (VPS). This approach achieves enterprise-grade scale at a fraction of the traditional cost.
This technical guide provides a comprehensive blueprint for architecting, deploying, and executing a distributed performance testing framework capable of simulating 50,000 concurrent users using k6 and Kubernetes on budget-friendly VPS infrastructure.
Architectural Overview: Distributed Load Generation
Simulating 50,000 concurrent users from a single machine is practically impossible due to hardware bottlenecks, specifically relating to CPU utilization, memory allocation, and network socket exhaustion. A single instance of a testing tool will hit the operating system's open file limits or saturate the network interface long before reaching the target scale. To overcome these physical limitations, we must adopt a distributed load generation architecture.
In a distributed architecture, the target load is segmented and distributed across multiple worker nodes (Load Generators). A centralized orchestrator coordinates the execution, aggregates real-time metrics, and ensures synchronized test cycles. Our self-hosted solution utilizes three core components:
- Virtual Private Servers (VPS): Low-cost, high-performance cloud compute instances acting as our physical or virtualized infrastructure layer.
- Kubernetes & k6 Operator: Kubernetes manages the lifecycle of our infrastructure, while the specialized k6 Operator automates the distribution, execution, and cleanup of the test scripts across the cluster.
- Grafana & Prometheus: A dedicated observability stack to collect, aggregate, and visualize performance metrics from both the load generators and the target application in real-time.
Key Architecture Principle: To accurately simulate 50,000 users, the load must be uniformly distributed across geographically optimized or network-isolated VPS instances to avoid creating artificial bottlenecks within the testing infrastructure itself.
Step 1: Provisioning and Optimizing the Budget VPS Cluster
The foundation of a cost-effective testing platform lies in selecting the right VPS provider and tuning the underlying operating system. Providers such as Hetzner, DigitalOcean, Linode, or Contabo offer excellent compute-to-cost ratios. For a 50,000 concurrent user test, we must estimate the required compute resources based on k6's memory blueprint.
Resource Estimation
Typically, a lean k6 virtual user (VU) script requires between 1MB and 5MB of memory, depending on the complexity of the JavaScript execution, header sizes, and response handling. Assuming an optimized script averaging 2MB per VU, simulating 50,000 VUs requires approximately 100GB of RAM across the entire cluster. To ensure adequate headroom for the OS and Kubernetes system daemons, we can architect the cluster using:
- 1x Master Node: 4 vCPU, 8GB RAM (Dedicated to cluster orchestration).
- 8x Worker Nodes: 4 vCPU, 16GB RAM each (Totaling 32 vCPU and 128GB RAM for load generation).
Operating System Kernel Tuning
Standard Linux VPS configurations are optimized for general-purpose workloads, not high-concurrency network throughput. Before deploying Kubernetes, you must modify the kernel parameters on all worker nodes to handle hundreds of thousands of concurrent network connections. Append the following configurations to /etc/sysctl.conf:
# Increase maximum number of open files and file descriptors
fs.file-max = 2097152
# Optimize ephemeral port range for outgoing connections
net.ipv4.ip_local_port_range = 1024 65535
# Enable fast reuse of TIME_WAIT sockets
net.ipv4.tcp_tw_reuse = 1
# Increase maximum backlog of connection requests
net.core.somaxconn = 32768
net.ipv4.tcp_max_syn_backlog = 16384
# Adjust TCP memory buffer allocations
net.ipv4.tcp_rmem = 4096 87380 16777216
et.ipv4.tcp_wmem = 4096 65536 16777216Apply these changes immediately using the command sudo sysctl -p. Failure to optimize these parameters will result in "connection reset by peer" or "socket allocation" errors during the initial ramp-up phase of your load test.
Step 2: Deploying a Lightweight Kubernetes Cluster
Running a standard upstream Kubernetes distribution can introduce unnecessary resource overhead on low-cost VPS instances. To maximize the resources available for load testing, we utilize K3s—a highly optimized, lightweight Kubernetes distribution developed by Rancher. K3s packages all components into a single binary under 100MB and reduces memory consumption significantly.
Master Node Installation
Execute the following command on your designated master node to initialize the control plane, ensuring that the embedded Let's Encrypt or default ingress controllers are disabled if you wish to conserve maximum resources:
curl -sfL [https://get.k3s.io](https://get.k3s.io) | sh -s - --disable traefikRetrieve the secure cluster token from the master node via cat /var/lib/rancher/k3s/server/node-token. This token is required to authenticate worker nodes joining the cluster.
Worker Node Registration
On each of the 8 worker nodes, execute the installation script while passing the master node's IP address and the secure token:
curl -sfL [https://get.k3s.io](https://get.k3s.io) | K3S_URL=https://:6443 K3S_TOKEN= sh - Verify the health and readiness of your cluster by running kubectl get nodes on the master node. You should see all nodes listed with a Ready status, establishing your distributed testing compute matrix.
Step 3: Installing the k6 Operator for Distributed Execution
The standard open-source version of k6 does not natively support multi-node clustering out of the box. To achieve distributed execution across our K3s cluster, we leverage the Grafana k6 Operator. This Custom Resource Definition (CRD) automates the process of partitioning a large-scale k6 test into smaller, independent test runs, deploying them across multiple Kubernetes Pods, and aggregating their execution state.
Install the k6 Operator onto your cluster via Helm by executing the following commands:
helm repo add grafana [https://grafana.github.io/helm-charts](https://grafana.github.io/helm-charts)
helm repo update
helm install k6-operator grafana/k6-operator --namespace k6-operator --create-namespaceThe operator monitors the cluster for custom K6 resources. When a new test configuration is detected, it automatically spins up a initializer pod to parse the test script, followed by multiple runner pods that execute the actual load generation across your worker nodes.
Step 4: Writing the High-Concurrency k6 Script
An optimized test script is critical when simulating 50,000 concurrent users. JavaScript execution within k6 uses an isolated ECMAScript 6 runtime environment for each Virtual User. To avoid excessive memory consumption, scripts must minimize heavy external libraries, avoid complex JSON manipulation inside the main loop, and use efficient threshold definitions.
Create a file named performance-test.js. This script targets a specific microservice endpoint, defining a ramp-up phase to smoothly transition to 50,000 VUs, maintaining that peak load, and then ramping down safely:
import http from 'k6/http';
import { check, sleep } from 'k6';
export const options = {
scenarios: {
distributed_load: {
executor: 'ramping-vus',
startVUs: 0,
stages: [
{ duration: '5m', target: 50000 }, // Smooth ramp-up to 50k VUs
{ duration: '10m', target: 50000 }, // Hold peak load for 10 minutes
{ duration: '3m', target: 0 }, // Gradual ramp-down
],
gracefulRampDown: '30s',
},
},
thresholds: {
http_req_failed: ['rate<0.01'], // Test fails if error rate exceeds 1%
http_req_duration: ['p(95)<500'], // 95% of requests must complete under 500ms
},
};
export default function () {
const url = '[https://api.yourtargetdomain.com/v1/resource](https://api.yourtargetdomain.com/v1/resource)';
const params = {
headers: {
'Content-Type': 'application/json',
'X-Test-Source': 'k6-distributed-vps',
},
};
const res = http.get(url, params);
check(res, {
'status is 200': (r) => r.status === 200,
'transaction verified': (r) => r.body.includes('success'),
});
// Introduce a dynamic pacing pause between 1 to 2 seconds
sleep(Math.random() * 1 + 1);
}Store this script as a Kubernetes ConfigMap so that the distributed runner pods can access it natively within the cluster environment:
kubectl create configmap k6-test-script --from-file=performance-test.jsStep 5: Executing the Distributed 50,000 VU Load Test
With our script safely stored within the cluster, we define a K6 Custom Resource manifest that explicitly instructs the k6 Operator how to distribute our workload. We will split the execution across 50 parallel runner instances. Each instance will dynamically inherit a slice of the overall load, executing exactly 1,000 concurrent users simultaneously.
Create a deployment file named k6-distributed-deployment.yaml:
apiVersion: k6.io/v1alpha1
kind: K6
metadata:
name: distributed-vps-load-test
namespace: default
spec:
parallelism: 50
script:
configMap:
name: k6-test-script
file: performance-test.js
runner:
resources:
limits:
cpu: "1000m"
memory: "2Gi"
requests:
cpu: "500m"
memory: "1Gi"
securityContext:
sysctls:
- name: net.core.somaxconn
value: "32768"
arguments: --tag test_run_id=vps-50k-benchmarkApply the manifest to initialize the test:
kubectl apply -f k6-distributed-deployment.yamlThe k6 Operator instantly parses this instruction, schedules the 50 runner pods across your 8 optimized VPS worker nodes, coordinates their network clocks, and begins generating synchronous traffic targeting your destination architecture. Monitor execution progress via command line logging: kubectl logs -f job/distributed-vps-load-test-initializer.
Step 6: Metric Aggregation and Observability
Running a distributed test generates vast quantities of performance metrics that must be structured and visualized. It is highly inefficient to analyze 50 individual log streams manually. To achieve real-time insight, configure your K6 resource definition to stream execution telemetry directly to a centralized Prometheus instance or Grafana Mimir data lake.
By modifying the arguments attribute within the custom manifest, you can instruct k6 to push real-time data using the native Prometheus Remote Write protocol:
arguments: --out experimental-prometheus-remote-write=url=http://prometheus-service:9090/api/v1/writeA pre-configured Grafana Dashboard (such as Dashboard ID 19665) visualizes critical performance metrics in real-time, focusing on:
- Request Rates (RPS): Total throughput managed by the target system.
- Error Rates: Percentages of HTTP 5xx, 4xx, or TCP connectivity dropping out under peak load.
- Latency Percentiles: Granular look at $p(50)$, $p(95)$, and $p(99)$ response intervals to identify long-tail distribution lag.
- Resource Consumption: Real-time CPU and memory metrics from your VPS workers to guarantee that the load generators did not experience saturation.
Conclusion: Enterprise Capabilities at Open-Source Costs
By leveraging open-source components like k6, K3s, and the k6 Operator, you can eliminate dependencies on costly commercial load testing platforms. Hosting a distributed testing infrastructure on an affordable VPS cluster provides engineering teams with complete architectural autonomy. It allows you to simulate extreme traffic events—such as 50,000 concurrent users—safely, repeatedly, and with deep observability.
With this infrastructure in place, your organization can proactively validate system resilience, optimize auto-scaling policies, and identify concurrency bottlenecks before they impact end-users, ensuring software stability at an optimal infrastructure price point.
