Back to articles
Technology Insight

Scaling Vector Search on Limited RAM: Optimizing pgvector with HNSW and Memory-Mapped Files

May 26, 2026

Introduction: The Memory Hurdle in Modern Vector Search

In the era of Generative AI and Retrieval-Augmented Generation (RAG), vector databases have transitioned from niche tools to infrastructure essentials. However, as organizations scale their applications, they often hit a common wall: the prohibitive cost of memory. High-dimensional embeddings, particularly when indexed using the Hierarchical Navigable Small World (HNSW) algorithm, are notoriously RAM-hungry.

For businesses utilizing Virtual Private Servers (VPS) to manage costs, the challenge is twofold. You need the low-latency performance of vector search, but you lack the enterprise-grade RAM capacity of dedicated clusters. This article explores a sophisticated technical strategy: leveraging pgvector and the underlying Linux Memory-Mapped Files (mmap) architecture to optimize large-scale vector search while maintaining a lean hardware footprint.

Understanding the HNSW Indexing Bottleneck

The HNSW algorithm is the gold standard for Approximate Nearest Neighbor (ANN) search due to its superior speed and recall. It functions by creating a multi-layered graph where each node represents a vector. Navigating these layers allows for $O(log(n))$ search complexity.

Why RAM Consumption Spikes

  • Graph Connectivity: Each vector requires additional storage for its adjacency lists (pointers to neighbors).
  • Vector Dimensions: A standard 1536-dimensional embedding (like OpenAI's text-embedding-3-small) consumes significant space per record.
  • In-Memory Requirement: Traditionally, HNSW performance relies on keeping the entire graph structure in RAM to avoid disk I/O latency.

When your index size exceeds your available VPS RAM, the system begins to swap, leading to a catastrophic drop in query performance. This is where tiered memory configuration becomes vital.

The Role of pgvector and PostgreSQL Buffer Cache

PostgreSQL, via the pgvector extension, manages data through its internal buffer cache. When you perform a search, the engine looks for the relevant index pages in memory. If they aren't there, it fetches them from the disk. In a standard configuration, this can be slow. However, by optimizing how PostgreSQL interacts with the OS Page Cache, we can achieve "near-RAM" speeds for datasets that are technically larger than our physical memory.

The Power of Memory-Mapped Files (mmap)

Linux utilizes mmap to map files or devices into memory. This allows applications to access file data as if it were in primary memory. In the context of pgvector, the operating system's kernel manages which parts of the HNSW index remain in RAM based on access patterns.

"By relying on the kernel's sophisticated Least Recently Used (LRU) algorithms, we can ensure that the 'hot' layers of our HNSW graph remain cached, while 'cold' data resides on high-speed NVMe storage."

Step-by-Step Strategy for VPS Optimization

1. Sizing Your Index Correctly

Before configuration, you must calculate the expected memory footprint. For an HNSW index in pgvector, the formula for memory usage is roughly:

$$Memory = Rows imes (Dimension imes 4 + M imes 8 imes 2)$$

Where M is the number of connections per element. Understanding this allows you to determine how much of your index will reside on disk versus RAM.

2. Tuning PostgreSQL for 'Hot' Data

To maximize the efficiency of memory-mapped files, you must balance your shared_buffers. Setting this too high can actually starve the OS Page Cache, which is more efficient at managing mmap-based index traversal.

  • shared_buffers: Set to 25% of total RAM.
  • effective_cache_size: Set to 75% of total RAM (informing the planner of available OS cache).
  • maintenance_work_mem: Increase this during index creation to speed up HNSW graph construction.

3. Leveraging NVMe for Tiered Storage

On a VPS, the latency of your storage medium is the ultimate fallback. Ensure your VPS uses NVMe SSDs. When the HNSW index traverses a node not in RAM, the fetch time from an NVMe drive is measured in microseconds, which—while slower than RAM—is often acceptable for sub-second search requirements in business applications.

Advanced Configuration: Prefaulting and Huge Pages

To further reduce the overhead of the memory-mapping system, consider the following Linux-level optimizations:

Enabling Transparent Huge Pages (THP)

Standard memory pages are 4KB. For large vector indices, the Translation Lookaside Buffer (TLB) can become a bottleneck. Using Huge Pages (typically 2MB) reduces the number of page table entries the CPU needs to track, improving the speed of memory access during graph traversal.

Index Warming

After a reboot or a service restart, your cache is "cold." Use a script to perform dummy searches or use the pg_prewarm extension to load your HNSW index into the buffer cache manually. This prevents the initial user queries from experiencing high latency.

Case Study: Scaling to 1 Million Vectors on 8GB RAM

Consider a scenario where a company needs to search 1,000,000 vectors of 768 dimensions. A pure in-memory HNSW index might require 12GB+ of RAM. On an 8GB VPS:

  1. The HNSW index is created with m=16 and ef_construction=64.
  2. The OS keeps the top layers of the HNSW graph (the most frequently accessed) in the 4GB of available page cache.
  3. The bottom layer (the largest portion) resides on NVMe.
  4. Result: Query latency increases from 5ms (full RAM) to 40ms (tiered memory), but the infrastructure cost is reduced by 60%.

Conclusion: Efficiency as a Competitive Advantage

Optimizing pgvector through memory-mapped files and strategic OS tuning transforms a VPS from a simple host into a high-performance vector engine. By understanding that not all data needs to be in RAM at all times, developers can scale their AI features sustainably.

As you implement these strategies, remember that the goal is not just raw speed, but the balance of cost, performance, and scalability. With HNSW and proper memory tiering, your PostgreSQL database becomes a formidable tool in your AI stack.

Scaling Vector Search on Limited RAM: Optimizing pgvector with HNSW and Memory-Mapped Files | DPTCloud