Scaling Beyond Redis: How to Build a Multi-VPS ARM Valkey Cluster for 1 Million RPS
Introduction: The Open-Source Caching Evolution
For over a decade, Redis stood as the undisputed backbone of high-performance web applications, powering session management, real-time analytics, and rapid caching layers. However, the tech landscape shifted dramatically in early 2024 when Redis transitioned from the open-source BSD license to a dual-license model (RSALv2 and SSPLv1). This pivot left enterprise architects and DevOps teams scrambling for a truly open-source alternative that maintained high compatibility and production-grade performance.
Enter Valkey. Backed by the Linux Foundation and supported by industry giants like AWS, Google Cloud, Oracle, and Ericsson, Valkey is a fully open-source, community-driven fork of Redis 7.2.4. It retains full drop-in compatibility while actively introducing performance improvements. In this comprehensive guide, we will explore how to architect and deploy a production-grade Valkey Cluster across a Multi-VPS ARM architecture, engineered specifically to handle a staggering 1 million requests per second (RPS) while optimizing your infrastructure budget.
Why ARM Architecture and Valkey are a Perfect Match
When engineering a data layer capable of handling 1 million RPS, infrastructure efficiency is paramount. Modern cloud computing has seen a massive shift toward ARM64-based processors (such as AWS Graviton, Ampere Altra, and OCI Ampere nodes). Deploying Valkey on ARM provides two distinct competitive advantages:
- Superior Price-to-Performance Ratio: ARM-based virtual private servers (VPS) typically cost 20% to 40% less than their x86 equivalents while delivering comparable, and sometimes superior, single-core memory bandwidth performance.
- Energy and Architectural Efficiency: Because Valkey is predominantly single-threaded for core data operations, it thrives on processors that offer deterministic, high-frequency per-core performance without the hyper-threading overhead common in x86 architectures.
By leveraging a Multi-VPS ARM setup, we distribute both memory capacity and network I/O across independent physical nodes, effectively bypassing the single-node hardware bottlenecks that typically stall massive scale-out strategies.
Architecting for 1 Million RPS: The Cluster Blueprint
To safely handle 1,000,000 requests per second with high availability, a standalone instance will not suffice. We must utilize a distributed Valkey Cluster. In a Valkey Cluster, the keyspace is divided into 16,384 logical hash slots. These slots are distributed across multiple primary nodes.
Production Rule of Thumb: For high availability and fault tolerance, a production cluster requires a minimum of three primary nodes and three replica nodes, spread across different physical infrastructure layers or availability zones.
To hit our target of 1 million RPS, our baseline architecture consists of:
- 6 Node Cluster: 3 Primary Nodes (handling read/write operations) and 3 Replica Nodes (handling read-scaling and failover).
- Server Specifications: Each VPS runs on 4 vCPUs (ARM64), 16GB RAM, with a dedicated 10 Gbps network interface.
- Load Distribution: Estimated throughput per ARM primary node is safely throttled around 350,000 RPS using optimized pipelining, aggregating to well over the 1,000,000 RPS milestone.
Step-by-Step Deployment Guide on Multi-VPS ARM
Let us walk through provisioning and configuring your high-performance Valkey Cluster across your multi-VPS ARM environment running Ubuntu 24.04 LTS.
Step 1: System Level Optimization
Before installing Valkey, the underlying Linux kernel must be tuned for high network throughput and low memory latency. Execute these configurations on all VPS nodes.
First, disable Transparent Huge Pages (THP), which can introduce significant memory latency spikes during background saving processes:
echo never > /sys/kernel/mm/transparent_hugepage/enabled
Next, optimize memory allocation behaviors and network backlogs by editing /etc/sysctl.conf:
# Ensure adequate memory overcommit for background snapshots
vm.overcommit_memory = 1
# Maximize network connection queues for massive traffic spikes
net.core.somaxconn = 65535
net.ipv4.tcp_max_syn_backlog = 65535
Apply the changes using sudo sysctl -p.
Step 2: Installing Valkey from Source on ARM64
Compiling from source ensures that Valkey leverages the exact vector instructions and optimizations unique to your ARM processor architecture.
- Install the necessary build dependencies:
sudo apt update && sudo apt install -y build-essential tcl-dev libjemalloc-dev git - Clone the official Valkey repository and checkout the latest stable release:
git clone [https://github.com/valkey-io/valkey.git](https://github.com/valkey-io/valkey.git) cd valkey git checkout 7.2 - Compile using jemalloc for superior memory management performance over standard glibc allocations:
make USE_JEMALLOC=yes sudo make install
Step 3: Configuring the Node Parameters
Create a dedicated configuration file at /etc/valkey/valkey-cluster.conf on each VPS. Below is a production-optimized blueprint for an ARM-based node:
# Network Configuration
port 6379
bind 0.0.0.0
protected-mode no
# Performance & Threading Tuning
protected-mode no
io-threads 4
io-threads-do-reads yes
maxmemory 12gb
maxmemory-policy volatile-lru
# Cluster Configurations
cluster-enabled yes
cluster-config-file nodes.conf
cluster-node-timeout 5000
# Persistence Configurations (Tuned for high throughput)
appendonly yes
appendfsync everysec
no-appendfsync-on-rewrite yes
Ensure that you adjust the maxmemory flag depending on your node's total RAM, leaving at least 25% of memory headroom for system processes and cluster synchronization buffers.
Step 4: Initializing the Valkey Cluster
Once the Valkey service is active on all separate VPS nodes, it is time to stitch them together into a unified cluster. Run the following command from your primary manager node, substituting the placeholders with your actual ARM VPS private IP addresses:
valkey-cli --cluster create \
10.0.1.10:6379 10.0.1.11:6379 10.0.1.12:6379 \
10.0.1.20:6379 10.0.1.21:6379 10.0.1.22:6379 \
--cluster-replicas 1
The system will prompt you with an allocation map displaying how the 16,384 hash slots are partitioned. Confirm by typing yes. Your highly distributed, 100% open-source cluster is now fully live.
Advanced Configuration tuning for 1M+ RPS
To push your multi-VPS cluster past the million-request milestone, standard out-of-the-box cluster orchestration isn’t quite enough. Implement these advanced modifications to unlock true architectural performance:
1. Activating I/O Threading
While the core execution of commands in Valkey is single-threaded, network handling (reading from and writing to sockets) can be heavily parallelized. By setting io-threads 4 in your configuration, you offload the network multiplexing to separate ARM cores, keeping the primary data execution thread completely unbottlenecked.
2. Client-Side Pipelining
Achieving 1 million RPS over a standard network without exhaustively choking network links requires query pipelining. Instead of executing a traditional atomic request-response flow for every single operation, application clients should utilize pipelining to batch dozens of commands into a single network packet. This exponentially reduces network context switching and system call friction across your VPS infrastructure.
Conclusion: A Future-Proof Data Layer
Migrating to a multi-VPS ARM architecture powered by Valkey represents a massive paradigm shift for modern infrastructure design. It proves that you do not need proprietary licenses or ultra-expensive x86 hardware configurations to execute complex, high-throughput, sub-millisecond workloads at enterprise scale.
By leveraging Valkey's robust open-source ecosystem alongside the structural and cost efficiencies of ARM processors, your organization can seamlessly achieve 1 million requests per second—ensuring your application's data layer remains fast, reliable, cost-effective, and fully free of vendor lock-in.
