Back to articles
Technology Insight

Scaling to 50,000 Concurrent Users: Self-Hosting Distributed Web Performance Testing with k6 and K3s on Budget VPS

May 26, 2026

Introduction: The Cost of Scale in Performance Testing

In modern software engineering, validating system resilience under peak load is a non-negotiable prerequisite for production deployment. However, engineering teams frequently encounter a significant bottleneck: the prohibitive cost of enterprise performance testing platforms. Simulating high-volume traffic—such as 50,000 concurrent virtual users (VUs)—via commercial SaaS solutions can quickly drain engineering budgets.

Fortunately, open-source tooling and cloud infrastructure efficiency have advanced dramatically. By combining Grafana k6, an ultra-efficient, developer-centric load testing tool, with K3s (a lightweight Kubernetes distribution), teams can self-host a highly scalable, distributed testing platform on budget-friendly Virtual Private Servers (VPS). This architectural blueprint guides you through setting up a distributed testing infrastructure capable of generating massive load without the commercial price tag.

---

Why k6 and K3s? The Architectural Synergy

Traditional load testing tools often suffer from high resource consumption or complex distribution mechanisms. Our chosen stack optimizes for both resource efficiency and operational simplicity:

  • Grafana k6: Written in Go, k6 utilizes a single OS thread per OS thread model for parallel execution, making it significantly more memory-efficient than Java-based alternatives like Apache JMeter. A single k6 process can comfortably manage thousands of VUs depending on script complexity.
  • K3s by Rancher: K3s is a highly optimized, fully compliant Kubernetes distribution designed for resource-constrained environments. It strips out legacy, alpha, and cloud-provider-specific plugins, reducing the memory footprint to under 512MB per node, making it ideal for budget VPS instances.
  • k6 Operator: This Kubernetes operator automates the orchestration of distributed k6 tests, managing the lifecycle of test initializers, runner pods, and test coordination flawlessly.
---

Infrastructure Blueprint: Sizing the Cluster

To simulate 50,000 concurrent VUs, we must plan our infrastructure to prevent the load generators themselves from becoming the bottleneck. Assuming a standard web test scenario (lightweight HTTP requests with JSON responses and a 1-second think time), a single vCPU can generally handle 1,000 to 2,500 VUs depending on network I/O and payload size.

For a conservative, resilient architecture, we recommend the following multi-node allocation:

Node TypeQuantityMinimum SpecificationsRole
Master Node (Control Plane)12 vCPU, 4GB RAMK3s Control Plane, k6 Operator, Monitoring Stack
Worker Nodes (Load Generators)54 vCPU, 8GB RAMDistributed k6 runner instances generating traffic
Pro-Tip: Select a VPS provider with unmetered internal bandwidth or high-capacity public networks (e.g., 1 Gbps to 10 Gbps port speeds) to ensure network interface cards (NICs) do not saturate during peak execution.
---

Step-by-Step Implementation Guide

Step 1: Deploying the K3s Cluster

First, initialize the K3s control plane on your designated master node. Execute the following deployment script via SSH:

curl -sfL [https://get.k3s.io](https://get.k3s.io) | sh -s - --disable traefik

Once initialized, extract the node token from /var/lib/rancher/k3s/server/node-token and use it to register your worker nodes:

curl -sfL [https://get.k3s.io](https://get.k3s.io) | K3S_URL=https://:6443 K3S_TOKEN= sh -

Step 2: Installing the k6 Kubernetes Operator

With your cluster operational, install the k6 Operator using Helm. This component interprets custom resources defined for distributed testing:

helm repo add k6-operator [https://grafana.github.io/helm-charts](https://grafana.github.io/helm-charts)
helm repo update
helm install k6-operator k6-operator/k6-operator

Step 3: Writing the Distributed k6 Test Script

Create a JavaScript file (test.js) configured to handle the target load. To maintain adaptability across different runner pods, use environment variables for execution parameters:

import http from 'k6/http';
import { check, sleep } from 'k6';

export const options = {
  discardResponseBodies: true,
  scenarios: {
    contacts: {
      executor: 'externally-controlled',
      duration: '10m',
      maxVUs: 10000, // Configured per runner pod
    },
  },
};

export default function () {
  const res = http.get('[https://your-target-system.com/api/v1/health](https://your-target-system.com/api/v1/health)');
  check(res, { 'status is 200': (r) => r.status === 200 });
  sleep(1);
}

Step 4: Creating the Custom Resource Definition (CRD)

To orchestrate the 50,000 VU load, bundle the script into a ConfigMap and define a K6 Custom Resource. We will distribute the total load across 5 runner pods, assigning 10,000 VUs to each:

apiVersion: k6.io/v1alpha1
kind: K6
metadata:
  name: distributed-perf-test
spec:
  parallelism: 5
  script:
    configMap:
      name: k6-test-script
      file: test.js
  arguments: --vus 10000

Deploy the ConfigMap and CRD using kubectl apply. The k6 Operator will automatically provision 5 separate runner pods across your worker nodes, synchronized to launch simultaneously.

---

Critical Optimization Strategies for Budget VPS

Running high-density load testing on affordable hardware requires fine-tuning the underlying Linux operating system. By default, standard Linux distributions limit resource allocations per process to safeguard stability.

1. Linux Kernel Tuning

Apply the following sysctl configurations to every worker node to handle massive concurrent socket connections and prevent the infamous "Too many open files" error:

fs.file-max = 500000
net.core.somaxconn = 65535
net.ipv4.tcp_max_syn_backlog = 65535
net.ipv4.ip_local_port_range = 1024 65535
net.ipv4.tcp_tw_reuse = 1

2. Resource Requests and Limits

Ensure your Kubernetes manifests strictly define resource constraints. If a k6 runner exceeds physical host limitations, the Linux Out-Of-Memory (OOM) killer will terminate the process, invalidating your test results. Align your pod limits to roughly 85% of total VPS host resources to leave overhead for OS and K3s background processes.

---

Observability: Monitoring Test Execution

A distributed performance test is only as good as the metrics it captures. To gain actionable insights, streaming metrics to a centralized dashboard is critical. The k6 Operator natively supports exporting telemetry to a Prometheus time-series database or an InfluxDB instance running on your master node.

By pairing this backend with a Grafana dashboard, you can track real-time system behaviors, including:

  • HTTP Request Rate (RPS): Total throughput achieved across all runners.
  • 95th and 99th Percentile Latencies: Precise customer-impact indicators, stripping out statistical anomalies.
  • HTTP Error Rate: The point at which the target application infrastructure begins to fail or drop packets.
---

Conclusion

Building a self-hosted distributed performance testing engine allows engineering organizations to achieve enterprise-level scale without enterprise-level pricing. By leveraging the low-overhead runtime of k6 and the minimal resource footprint of K3s, a collection of budget VPS instances can seamlessly simulate 50,000 concurrent visitors.

This framework empowers development teams to run load tests frequently within CI/CD pipelines, ensuring infrastructure bottlenecks are exposed and resolved long before they impact end-users in production environments.

Scaling to 50,000 Concurrent Users: Self-Hosting Distributed Web Performance Testing with k6 and K3s on Budget VPS | DPTCloud