Back to articles
Technology Insight

Scaling to Millions: Advanced Linux Kernel Tuning for High-Concurrency WebSocket VPS Servers

June 3, 2026

Introduction: The Architectural Challenge of Real-Time WebSockets

In the modern enterprise ecosystem, real-time data delivery is no longer a luxury—it is a core operational requirement. Applications ranging from financial trading platforms and collaborative SaaS tools to live gaming and IoT telemetry rely heavily on the WebSocket protocol. Unlike traditional HTTP requests that follow a short-lived, stateless request-response lifecycle, WebSockets establish persistent, bidirectional, and stateful TCP connections.

When scaling these real-time applications to support millions of concurrent connections (C10M challenge) on Virtual Private Servers (VPS), standard Linux kernel configurations fail rapidly. The default settings of most Linux distributions are optimized for general-purpose workloads, not high-density, long-lived TCP connections. Without precise system-level modifications, production environments face critical bottlenecks, including packet drops, high latency, connection timeouts, and memory exhaustion. This technical guide outlines the definitive Linux kernel optimization strategy required to transform a standard VPS into a high-throughput, low-latency WebSocket powerhouse.

1. Breaking the Limits: File Descriptors and System Resources

In Unix-like operating systems, "everything is a file." Every incoming WebSocket connection corresponds to a network socket, which requires an allocated file descriptor (FD). By default, Linux imposes restrictive limits on the number of open file descriptors to prevent a single process from monopolizing system resources.

System-Wide Configuration

To support millions of simultaneous connections, the system-wide limit for open files must be dramatically increased. This is managed via the kernel parameter fs.file-max. To view the current limit, check the virtual file system:

cat /proc/sys/fs/file-max

To persistently scale this limit across reboots, edit the /etc/sysctl.conf file and add the following entries:

fs.file-max = 20971520
fs.nr_open = 20971520

Here, we set the maximum limit to over 20 million files, providing a safe ceiling for both the sockets and the application-level files required during operations.

User-Level and Process-Level Security Limits

Modifying system-wide parameters is insufficient if the user running the application binary is constrained by lower thresholds. System resource usage is controlled via /etc/security/limits.conf. For an enterprise application runner (e.g., a dedicated system user named websocket_user), configure both soft and hard limits:

websocket_user soft nofile 10485760
websocket_user hard nofile 10485760
websocket_user soft nproc  65536
websocket_user hard nproc  65536

The nofile directive scales open file descriptors, while nproc increases the maximum number of processes/threads the user can spawn, preventing thread starvation under high compute loads.

2. Redefining the Linux Network Stack (sysctl.conf Optimization)

The Linux network stack requires granular tuning to handle massive volumes of incoming TCP handshake requests and maintain stability for idle connections. The sysctl utility is used to modify kernel parameters at runtime.

Managing the Connection Backlog

When millions of clients attempt to connect simultaneously (e.g., during a service recovery or peak traffic event), the kernel's connection queues can overflow, leading to dropped connection attempts. Optimize the backlog using three critical parameters:

  • net.core.somaxconn: The maximum number of backlogged sockets waiting for an accept() call. Elevating this prevents connection drops during connection bursts.
  • net.ipv4.tcp_max_syn_backlog: The maximum number of half-open connections that the system can track in the SYN queue before establishing a full TCP handshake.
  • net.core.netdev_max_backlog: Increases the number of packets allowed to queue on the network interface card (NIC) input side before being processed by the CPU.

Apply these values in /etc/sysctl.conf:

net.core.somaxconn = 65535
net.ipv4.tcp_max_syn_backlog = 65535
net.core.netdev_max_backlog = 100000

Optimizing TCP Timers and Reusing Sockets

WebSocket connections are stateful, but when clients disconnect, sockets enter transient states like TIME_WAIT. If millions of connections close rapidly, the system will run out of ephemeral ports. To prevent port exhaustion, configure the kernel to aggressively recycle and reuse sockets:

net.ipv4.tcp_tw_reuse = 1
net.ipv4.tcp_fin_timeout = 15
net.ipv4.ip_local_port_range = 1024 65535

By shortening the tcp_fin_timeout from the default 60 seconds to 15 seconds, resources are freed significantly faster. Expanding the ip_local_port_range maximizes the outbound ephemeral ports available per local IP address.

3. Memory Optimization for High-Density TCP Buffers

Memory is typically the primary limiting factor when scaling to millions of concurrent WebSockets. Each TCP socket requires allocated memory for its read and write buffers. If a single connection consumes 64 KB of memory, 1 million connections will require roughly 64 GB of RAM purely for network buffers, excluding the application overhead.

Granular Auto-Tuning of TCP Memory Buffers

Linux features an automatic TCP buffer tuning mechanism. We must configure the minimum, default, and maximum memory values (measured in bytes) to keep the footprint low for idle connections while allowing active connections to scale dynamic buffer sizes when transferring large payloads.

net.ipv4.tcp_rmem = 4096 8192 16777216
net.ipv4.tcp_wmem = 4096 8192 16777216
net.core.rmem_max = 16777216
net.core.wmem_max = 16777216

In this configuration, the default buffer size is set aggressively low at 8 KB. For 1 million concurrent connections, this restricts the base memory consumption to approximately 8 GB, making high-density concurrency achievable on cost-effective enterprise VPS nodes. When data spikes occur, individual sockets can dynamically scale up to 16 MB as dictated by network constraints.

4. Preventing Silences: Advanced TCP Keepalive Tuning

WebSocket connections can pass through various intermediate network infrastructure components, including firewalls, load balancers, and NAT gateways. Many of these stateful network devices drop inactive connections from their routing tables without notifying the server or client, creating "zombie" sockets.

To keep connections alive and accurately detect dead clients, adjust the global TCP keepalive probes:

net.ipv4.tcp_keepalive_time = 300
net.ipv4.tcp_keepalive_intvl = 15
net.ipv4.tcp_keepalive_probes = 5

This ensures that the kernel transmits a keepalive probe after 300 seconds (5 minutes) of inactivity, repeats the probe every 15 seconds if unacknowledged, and forcefully terminates the connection after 5 consecutive failed attempts. This frees up critical system memory from abandoned client connections.

5. Maximizing VPS CPU Efficiency: Interrupt Handling and Epoll

High-concurrency networking creates massive CPU overhead due to hardware and software interrupts generated by network interface cards (NICs). To ensure sub-millisecond real-time performance, optimal processor utilization is mandatory.

Leveraging the Epoll Event Loop

Modern high-performance runtimes (such as Node.js, Go, or Netty/Java) rely heavily on the Linux epoll (Event Poll) mechanism instead of deprecated subsystems like select() or poll(). Epoll provides an $O(1)$ time complexity, ensuring that system performance scales linearly regardless of whether 10 or 10,000,000 file descriptors are actively monitored. Verify that your application architecture natively abstracts epoll for asynchronous non-blocking I/O operations.

NIC Interrupt Distribution (IRQ Balance)

In many VPS environments, network interrupts are routed exclusively to CPU Core 0, creating a processing bottleneck while remaining cores sit idle. Ensure that the irqbalance daemon is enabled and running on your VPS distribution to dynamically distribute hardware interrupt requests across all available vCPUs:

systemctl enable --now irqbalance

Conclusion: The Deployment Checklist

Transforming a Linux VPS into an infrastructure capable of sustaining millions of concurrent WebSocket connections requires a systematic approach to kernel level optimization. To apply all changes permanently, execute the following command after updating your /etc/sysctl.conf file:

sysctl -p

By strictly managing open file descriptors, fine-tuning the TCP network layer backlogs, minimizing memory allocations per socket, and implementing aggressive zombie socket reclamation, your infrastructure will deliver superior real-time performance. Always execute load testing using synthetic traffic generators like tsung or wrk2 in a staging environment to validate your configurations before pushing to production live environments.

Scaling to Millions: Advanced Linux Kernel Tuning for High-Concurrency WebSocket VPS Servers | DPTCloud