Back to articles
Technology Insight

Scaling Large-Scale Vector Search on a Budget: Optimizing pgvector with Memory-Mapped Files and HNSW on VPS

May 26, 2026

Introduction: The Cost Crisis of Large-Scale Vector Search

In the era of Generative AI and Large Language Models (LLMs), vector search has become the backbone of modern data retrieval. From retrieval-augmented generation (RAG) to recommendation systems, high-dimensional vector embeddings are everywhere. However, engineering teams face a stark reality: vector search is notoriously memory-intensive.

When utilizing PostgreSQL's powerful pgvector extension with the Hierarchical Navigable Small World (HNSW) index, the golden rule has always been to keep the entire index resident in RAM to avoid devastating latency penalties. But what happens when your vector dataset grows into tens of gigabytes, or even terabytes, while your budget confines you to a virtual private server (VPS) with limited RAM? This technical guide explores how to build a highly optimized, tier-memorized architecture on a VPS, leveraging Memory-Mapped Files (mmap) via Linux virtual memory management to scale HNSW indexes far beyond physical RAM boundaries.

The Core Challenge: HNSW Index Memory Footprint

To understand the optimization strategy, we must first look at how the HNSW index operates within pgvector. HNSW constructs a multi-layer graph structure to achieve logarithmic search complexity. Each vector node maintains lists of bidirectional links to its nearest neighbors across multiple layers.

The memory consumption of an HNSW index in pgvector can be calculated using the following variables:

  • d: The dimension of the vector (e.g., 1536 for OpenAI's text-embedding-3-large or 768 for standard BERT models).
  • m: The maximum number of bidirectional links per node (configured via m parameter).
  • Size per vector: (d * 4 bytes) + (m * 2 * 4 bytes) + overhead.

For a dataset of 10 million vectors with 1536 dimensions, the raw vectors alone require roughly 61 GB of storage. When adding the HNSW graph overhead, the required memory quickly balloons to 80+ GB. On a standard VPS, provisioning 96 GB of RAM can be prohibitively expensive. If the system runs out of physical memory, the OS kernel's Out-Of-Memory (OOM) killer will instantly terminate the PostgreSQL process.

The Architectural Solution: Tiered Memory & Memory-Mapped Files

To bypass physical RAM limitations without sacrificing system stability, we can design a tiered memory architecture. Instead of forcing the entire HNSW index to reside in volatile RAM, we allow the operating system to treat high-speed NVMe SSD storage as an extended memory tier through Memory-Mapped Files (mmap) and aggressive virtual memory swapping.

PostgreSQL inherently relies on the OS page cache. When pgvector queries the HNSW index, it accesses data via shared buffers. By correctly configuring the Linux virtual memory subsystem, we can map the large index files into the process's virtual address space. The system treats the high-speed NVMe SSD as "Layer 2 Memory." When a specific node in the HNSW graph is required during a graph traversal, the OS fetches that specific disk page into RAM. If the page isn't needed immediately, it is seamlessly evicted back to disk.

Step-by-Step Guide to Configuring a VPS for Large-Scale pgvector

1. Optimizing the Operating System Layer

First, we must prepare the Linux kernel to handle heavy, random read operations on swap memory efficiently. Standard swap configurations are optimized for idle data eviction, which is lethal for real-time vector search. We need to alter this behavior.

First, provision a dedicated swap file on a high-performance NVMe SSD. Avoid standard cloud HDD blocks at all costs, as random read latencies will render your search times unusable.

# Create a 64GB swapfile for a 16GB/32GB RAM VPS
sudo fallocate -l 64G /swapfile
sudo chmod 600 /swapfile
sudo mkswap /swapfile
sudo swapon /swapfile

# Make it persistent
echo '/swapfile none swap sw 0 0' | sudo tee -a /etc/fstab

Next, adjust the kernel's memory management behavior by modifying /etc/sysctl.conf. We want to configure vm.swappiness carefully. While a low swappiness (e.g., 10) prevents normal apps from swapping, we want a moderate value (e.g., 20 to 30) combined with optimized file caching preferences so the kernel understands it should cache index pages aggressively:

# Reduce aggressive swapping but allow it for large memory mappings
vm.swappiness = 30
# Increase maximum memory maps allowed for large processes
vm.max_map_count = 262144
# Optimize page cache eviction behavior
vm.vfs_cache_pressure = 50

Apply the changes instantly using sudo sysctl -p.

2. Tuning PostgreSQL and pgvector Configurations

PostgreSQL must be explicitly informed that it does not have the luxury of holding everything in its dedicated shared_buffers. We must deliberately set shared_buffers to a value lower than the total index size, forcing reliance on the OS page cache and mmap mechanisms, while keeping enough headroom for active query processing.

Modify your postgresql.conf file with the following directional guidelines:

Recommended Configuration for a 32GB RAM VPS handling an 80GB Index:
shared_buffers = 8GB (25% of physical RAM)
work_mem = 64MB (To allow fast sorting and graph building per connection)
maintenance_work_mem = 4GB (Crucial for building the HNSW index initially)
effective_cache_size = 24GB (Tells the planner how much total memory + OS cache is available)

By keeping shared_buffers at 8GB, we leave 24GB of physical RAM available entirely for the operating system to act as a dynamic sliding-window cache for the HNSW index files mapped from the NVMe storage.

3. Constructing the HNSW Index Optimally

When creating the HNSW index via pgvector, the parameters m and ef_construction play a critical role in both precision and ultimate memory size. Lowering m reduces the memory footprint significantly, which translates directly to fewer random disk I/O operations when reading pages via mmap.

-- Creating an optimized HNSW index for memory-constrained environments
CREATE INDEX ON items USING hnsw (embedding vector_cosine_ops)
WITH (m = 16, ef_construction = 64);

While m = 16 provides a slightly less dense graph than the default m = 48, it dramatically slashes the index size by over 50%, ensuring a much higher percentage of the active graph layers can fit entirely inside the physical OS page cache.

Benchmarking Performance and Mitigating Latency

What can you expect regarding performance when shifting from a fully RAM-resident index to an mmap/SSD-tiered architecture?

  • Pure RAM Execution: Query latencies typically range between 2ms to 10ms.
  • Memory-Mapped NVMe SSD Execution: Query latencies typically scale to 15ms to 45ms.

While a 3x to 5x increase in latency might sound daunting on paper, in many real-world business scenarios (such as asynchronous background processing, document classification, or internal enterprise search), a 30ms response time is entirely acceptable—especially when it saves thousands of dollars per month in infrastructure overhead.

To maintain peak performance in this tiered setup, it is vital to execute a warm-up strategy after a server reboot or database restart. This forces the OS to map and read the index files sequentially into memory before users start querying:

-- Force a sequential scan over the index to pre-warm the OS page cache
SELECT COUNT(*) FROM items WHERE embedding <=> '[0.1, 0.2, ...]'::vector < 1.0;

Conclusion: Democratizing High-Scale Vector Search

You do not always need enterprise-grade, massive-RAM cloud instances to deliver cutting-edge semantic search features. By coupling PostgreSQL's pgvector with deep Linux virtual memory tuning, you can transform a standard, cost-effective VPS into a powerful vector processing machine. By shifting the paradigm from "RAM-only" to "Tiered NVMe-Mapped Memory," you strike an optimal balance between infrastructure costs and system performance, ensuring your AI applications can scale sustainably.

Scaling Large-Scale Vector Search on a Budget: Optimizing pgvector with Memory-Mapped Files and HNSW on VPS | DPTCloud