Back to articles
Technology Insight

Scaling Dragonfly: Reaching 1 Million OPS on NVMe VPS Infrastructure

May 27, 2026

Introduction: The Quest for 1 Million OPS on Commodity Infrastructure

In the era of high-throughput web applications, real-time analytics, and instantaneous caching, database performance is often the ultimate bottleneck. For years, Redis has been the go-to in-memory key-value store. However, as horizontal scaling demands increase and hardware capabilities evolve, standard single-threaded architectures face strict physical boundaries. Enter Dragonfly—a modern, ultra-fast in-memory data store designed specifically to leverage multi-threaded architectures and modern NVMe hardware.

Achieving 1 million Operations Per Second (OPS) used to require massive, expensive dedicated bare-metal servers. Today, with the maturity of NVMe-powered Virtual Private Servers (VPS) and Dragonfly's advanced shared-nothing architecture, this milestone is highly attainable on a budget. This comprehensive guide outlines the exact technical configurations, optimization strategies, and architectural blueprints required to build, scale, and benchmark a Dragonfly cluster capable of sustaining 1M+ OPS.

1. Understanding Dragonfly’s Core Architecture

To optimize Dragonfly effectively, one must understand why it fundamentally outperforms traditional key-value stores like Redis. While Redis relies on a single-threaded event loop (requiring multi-process cluster topologies to scale vertically), Dragonfly utilizes a shared-nothing multi-threaded architecture.

  • VLL (Virtual Lock Manager): Dragonfly uses an innovative locking mechanism that allows multi-threaded execution over a single shared memory space without the classic contention overhead found in traditional multi-threaded databases.
  • Thread-per-Core Model: Dragonfly binds exactly one execution thread to each CPU core. This eliminates CPU context switching, maximizes L1/L2 cache locality, and ensures linear performance scaling as you add CPU cores.
  • Advanced Memory Efficiency: Dragonfly implements novel data structures like dashmap, which drastically reduce memory serialization overhead and memory fragmentation compared to Redis dicts.

When running on a modern VPS, this means Dragonfly can fully utilize all vCPUs allocated to the instance seamlessly out of the box, making it exceptionally well-suited for high-density cloud environments.

2. Selecting and Preparing the NVMe VPS Infrastructure

Not all VPS configurations are created equal. To hit the 1 million OPS threshold, your underlying compute, storage, and networking layers must be fine-tuned to eliminate hypervisor-induced bottlenecks.

Hardware Prerequisites

For a production-grade Dragonfly cluster capable of 1M OPS, we recommend distributing the load across a 3-node cluster. Each node should meet the following specifications:

  • Compute: At least 8 vCPUs (Compute-Optimized instances with high single-core clock speeds, such as AMD EPYC or Intel Xeon scalable processors).
  • Memory: 32GB RAM (Ensure high memory speed and sufficient capacity for your dataset).
  • Storage: Local NVMe SSDs. Standard SATA SSDs or network-attached block storage (like AWS EBS or DigitalOcean Block Storage) introduce unacceptable I/O wait times during snapshotting and cache eviction.
  • Network: Minimal 10 Gbps network interface with low-latency private networking enabled between nodes.

3. Operating System & Kernel Optimizations

Before launching Dragonfly, the Linux kernel must be modified to handle hundreds of thousands of concurrent network connections and high-throughput memory allocations.

Note: Standard Linux distributions are tuned for general-purpose workloads, not high-frequency low-latency data tiers. Skipping kernel tuning will result in connection drops and latency spikes long before reaching 1M OPS.

Network Stack Tuning (sysctl.conf)

Append the following parameters to your /etc/sysctl.conf file to optimize the TCP/IP stack for extreme throughput:

# Maximize concurrent connection backlog
net.core.somaxconn = 65535
net.ipv4.tcp_max_syn_backlog = 65535

# Optimize TCP window sizes and buffer memory allocations
net.core.rmem_max = 16777216
net.core.wmem_max = 16777216
net.ipv4.tcp_rmem = 4096 87380 16777216
net.ipv4.tcp_wmem = 4096 65536 16777216

# Enable fast recycling of TCP connections
net.ipv4.tcp_tw_reuse = 1
net.ipv4.tcp_fin_timeout = 15

Apply changes immediately using sudo sysctl -p.

Memory Management & Huge Pages

Unlike Redis, which often benefits from disabling Transparent Huge Pages (THP) due to fork-based snapshotting overhead, Dragonfly handles snapshotting using an asynchronous serialization technique. Therefore, keeping Transparent Huge Pages (THP) enabled or explicitly allocating Explicit Huge Pages can significantly improve memory lookup speeds and reduce Translation Lookaside Buffer (TLB) misses.

Ensure your system limits allow for sufficient open files by modifying /etc/security/limits.conf:

* soft nofile 100000
* hard nofile 100000

4. Configuring Dragonfly for Peak Performance

Dragonfly configuration is streamlined via command-line flags. To maximize efficiency on an NVMe VPS, specific memory management and I/O arguments must be passed to the daemon.

An optimized production execution command looks like this:

dragonfly --proactor_threads=8 \
          --mem_oom_limit_mb=28000 \
          --dbfilename=/mnt/nvme/dragonfly.snapshot \
          --cache_mode=true \
          --keys_output_limit=1000 \
          --admin_port=9000

Key Flag Breakdown:

  1. --proactor_threads: Matches the exact number of vCPUs. This enforces the thread-per-core model, pinning processing loops to dedicated hardware threads.
  2. --mem_oom_limit_mb: Restricts Dragonfly's memory consumption slightly below the system total (e.g., 28GB out of 32GB) to leave headroom for OS processes and prevent kernel OOM killer interventions.
  3. --cache_mode: When set to true, Dragonfly acts as a smart cache, automatically evicting least-recently-used (LRU) items when memory limits are reached, maintaining high availability without crashing.
  4. --dbfilename: Explicitly points to the high-performance local NVMe mount point, ensuring point-in-time snapshots do not block memory operations.

5. Building a Scalable Distributed Cluster

While a single well-specced Dragonfly instance can achieve massive throughput, scaling across a distributed cluster ensures high availability, redundancy, and load distribution required for true enterprise applications.

Dragonfly seamlessly supports the standard Redis Cluster protocol, allowing existing application clients (like Predis, Jedis, or ioredis) to interact with it without code rewrites. By configuring a multi-node cluster with replication, read operations can be scaled horizontally across replicas, while write operations are channeled to master nodes.

Utilizing Dragonfly’s native replication flags (--masterauth and --replicaof), data synchronization over the internal private 10 Gbps network happens asynchronously, ensuring master execution loops spend zero time waiting for replication acknowledgments.

6. Benchmarking and Verifying 1 Million OPS

To accurately verify that your optimized configuration can sustain 1,000,000 operations per second, you must benchmark using a dedicated client machine to avoid resource contention on the database nodes. We use redis-benchmark or Dragonfly's native emulation utilities from an external VPS within the same private network.

Execute the following high-concurrency test from the benchmark client:

redis-benchmark -h  -p 6379 -c 500 -n 10000000 -d 256 -t set,get -P 16 --cluster

Analyzing the Parameters:

  • -c 500: Simulates 500 concurrent client connections mimicking a high-traffic microservices mesh.
  • -n 10000000: Sends a total of 10 million requests to gather statistically significant average and tail latency figures.
  • -d 256: Uses a 256-byte payload size, typical for authentication tokens, session data, or cached API payloads.
  • -P 16: Uses a pipeline depth of 16. Pipelining groups multiple commands into a single network packet, drastically reducing network round-trip time (RTT) overhead and maximizing Dragonfly's internal processing pipeline.

Upon execution, monitoring your instances will show CPU utilization evenly distributed across all configured proactor threads, with latency profiles remaining strictly sub-millisecond even as total throughput scales past the 1 million OPS threshold.

Conclusion

Scaling a key-value database to 1 million OPS no longer requires complex, multi-million dollar infrastructure. By shifting from legacy architectures to Dragonfly and backing it with fast NVMe VPS instances, businesses can unlock elite-tier performance at a fraction of the traditional cost. Through proper kernel network layer modifications, precise thread mapping, and structured pipelining, your infrastructure will easily handle the most demanding modern workloads with bulletproof stability.

Scaling Dragonfly: Reaching 1 Million OPS on NVMe VPS Infrastructure | DPTCloud