Back to articles
Technology Insight

Scaling to Millions: How to Optimize Your VPS for 1 Million Concurrent WebSocket Connections

May 27, 2026

Introduction to the 1-Million Connection Challenge

In the era of real-time web applications, features like live chat, financial tickers, collaborative tools, and multiplayer gaming have shifted from luxuries to standard expectations. WebSockets have emerged as the industry standard protocol for delivering these low-latency, bidirectional communication channels. However, transitioning from a few thousand connections to supporting 1 million concurrent WebSocket connections on a Virtual Private Server (VPS) introduces profound engineering challenges.

At this scale, the primary bottleneck is rarely the CPU; instead, the system hits hard limits regarding network I/O, memory consumption, and operating system configuration limits. By default, standard Linux distributions are tuned for general-purpose server workloads, meaning they will actively block or crash under the weight of a million open sockets. This comprehensive guide details the precise system-level optimizations, network architecture adjustments, and application-level strategies required to unlock maximum concurrency from your VPS infrastructure.

---

Understanding the Linux Network Stack Bottlenecks

To optimize a system, one must first understand what limits it. When a client establishes a WebSocket connection, it starts as a standard HTTP request and upgrades to a persistent TCP connection. Maintaining 1 million persistent connections exposes three critical operating system constraints:

  • File Descriptor Limits: In Linux, "everything is a file." Every open network socket requires a file descriptor (FD). Default configurations typically cap FDs at 1,024 per process—miles away from our target.
  • The Epoll Mechanism: Traditional I/O multiplexing methods like select or poll degrade linearly ($O(N)$) as connections grow. Modern systems must rely on epoll ($O(1)$) to handle state changes efficiently across millions of active links.
  • TCP Memory Footprint: Each TCP connection allocates a buffer for sending and receiving data. If every connection consumes even 32 KB of RAM, 1 million connections will require at least 32 GB of memory just for network buffers, excluding application overhead.
---

Step 1: Kernel Tuning and System Limits (sysctl.conf)

The first line of defense is modifying the Linux kernel parameters via the /etc/sysctl.conf file. These settings alter how the kernel manages networking tables, timeouts, and memory allocation.

Increasing File Descriptors Globally

Before configuring application-specific limits, you must instruct the operating system to allow a massive number of concurrent open files. Append the following lines to your system configuration:

fs.file-max = 2097152
fs.nr_open = 2097152

Note: We set the limit to roughly 2 million to ensure ample headroom for system processes alongside our 1 million application connections.

Optimizing TCP Window and Buffer Sizes

To minimize the memory footprint per connection, we must aggressively tune the TCP read and write memory vectors. While large buffers maximize throughput for file transfers, WebSocket messages are typically small, allowing us to reduce buffer sizes significantly:

net.ipv4.tcp_rmem = 4096 87380 16777216
net.ipv4.tcp_wmem = 4096 65536 16777216

Additionally, adjust the global TCP memory boundaries, which tell the kernel when to start rationing memory or enforcing pressure tactics (measured in pages, usually 4KB):

net.ipv4.tcp_mem = 786432 1048576 2621440

Accelerating Connection Queues

When millions of users attempt to connect simultaneously (e.g., during a service reconnection wave), the connection backlog queues fill instantly. Prevent connection drops by increasing the maximum backlog sizes:

net.core.somaxconn = 65535
net.ipv4.tcp_max_syn_backlog = 65535
net.core.netdev_max_backlog = 65535
Pro-Tip: Enable TCP SYN Cookies (net.ipv4.tcp_syncookies = 1) to protect your VPS from SYN flood DDoS attacks during high connection spikes, ensuring legitimate WebSocket handshakes still go through.
---

Step 2: Modifying Security Limits (limits.conf)

Even if the kernel is capable of handling millions of files, Linux enforces per-user and per-process restrictions to prevent resource starvation. To alter these, open /etc/security/limits.conf and define explicit soft and hard limits for the user account running your WebSocket application or reverse proxy:

* soft nofile 1048576
* hard nofile 1048576
root soft nofile 1048576
root hard nofile 1048576

To ensure these limits take effect immediately upon user session initialization, verify that your PAM configuration (/etc/pam.d/common-session) includes the directive: session required pam_limits.so.

---

Step 3: Port Exhaustion and Local IP Binding

A standard networking constraint often surprises engineers trying to scale WebSockets: ephemeral port exhaustion. By default, an outbound or inbound proxy connection via a single IP address is bound by the available local port range (typically 32768 to 60999, providing roughly 28,000 ports).

To expand this pool, expand the range to its absolute maximum constraints:

net.ipv4.ip_local_port_range = 1024 65535

Even with this expansion, a single IP can only yield about 64,000 connections per destination IP. To achieve 1 million connections, you must assign multiple virtual IP addresses to your VPS network interface card (NIC). By binding your reverse proxy to handle requests across 16 or 32 distinct IP addresses, the available tracking combinations ($Source\,IP \times Source\,Port \times Destination\,IP \times Destination\,Port$) multiply, easily accommodating the 1-million connection threshold.

---

Step 4: Optimizing Reverse Proxies (Nginx & HAProxy)

Placing an optimized reverse proxy in front of your application server is essential for SSL termination, load balancing, and traffic scrubbing. Here is how to construct your configurations for extreme concurrency.

Nginx Configuration Tuning

In your nginx.conf, maximize the worker connections and switch the event model explicitly to epoll:

worker_processes auto;
worker_rlimit_nofile 1048576;

events {
    use epoll;
    worker_connections 500000;
    multi_accept on;
}

Within the http block, ensure that timeouts are prolonged to prevent unneeded closures of idle connections, and keepalive configurations are aggressively tuned:

http {
    keepalive_timeout 65;
    keepalive_requests 100000;
    
    upstream websocket_backend {
        server 127.0.0.1:8080;
        keepalive 1000; # Keep connections alive to backend application
    }
}
---

Step 5: Application-Level Architecture Strategies

System tweaks mean nothing if the code running your WebSocket server is inefficient. At 1 million connections, language choices and memory management patterns dictate success.

1. Choose Concurrent, Non-Blocking Runtimes

Thread-per-connection models (like traditional Apache or standard Ruby/Python configurations) will fail instantly. Instead, rely on highly asynchronous, event-driven architectures:

  • Go (Golang): Highly recommended due to lightweight goroutines that consume as little as 2 KB of memory initially.
  • Node.js: Known for its single-threaded event loop, but must be run in a cluster mode utilizing all CPU cores.
  • Rust: Excellent for raw performance and predictable memory management without garbage collection pauses.

2. Manage the Garbage Collector (GC)

In languages like Go or Node.js, allocations triggered by incoming messages accumulate quickly. When the Garbage Collector runs to clean up memory, it can cause "Stop-the-World" pauses. If your server pauses for just 1 second while managing 1 million active nodes, thousands of connections will timeout and drop simultaneously, creating a catastrophic re-connection storm. Use object pooling (e.g., sync.Pool in Go) to reuse byte buffers rather than allocating fresh memory allocations for every incoming WebSocket frame.

---

Conclusion and Monitoring Requirements

Achieving 1 million concurrent WebSocket connections on a VPS is an engineering milestone that transforms your real-time infrastructure capacity. By shifting from default operating system assumptions to a highly tailored environment, you ensure your server safely utilizes every byte of RAM and CPU cycle available.

Before launching to production, validate your environment using open-source load-testing frameworks like K6, Tsung, or Artillery configured across separate distributed testing instances. Coupled with real-time monitoring via Prometheus and Grafana to track metrics like open FDs, memory allocation, and TCP retransmissions, your system will remain secure, stable, and incredibly responsive under massive loads.