Back to articles
Technology Insight

Scaling Large-Scale Vector Search on a Budget: Optimizing VPS with Memory-Mapped Files and HNSW in pgvector

May 25, 2026

Introduction: The Memory Crisis in Large-Scale Vector Search

In the era of Generative AI and Large Language Models (LLMs), production applications increasingly rely on vector embeddings for semantic search, retrieval-augmented generation (RAG), and recommendation engines. However, engineering teams scaling these systems quickly hit a formidable bottleneck: memory consumption.

The Hierarchical Navigable Small World (HNSW) algorithm is the gold standard for Approximate Nearest Neighbor (NN) search due to its exceptional query speed and high recall. Yet, HNSW is notoriously RAM-hungry. In a standard PostgreSQL deployment using the pgvector extension, the entire HNSW graph index typically needs to reside in RAM to ensure high-performance queries. For millions of high-dimensional vectors (such as OpenAI's 1536-dimensional text-embedding-3-small), the hardware costs can escalate rapidly, making standard Virtual Private Servers (VPS) look insufficient.

But what if you could scale vector search on a budget without compromising precision? By leveraging Memory-Mapped Files (mmap) alongside structural optimizations within pgvector, you can tier your memory architecture. This allows your operating system to seamlessly stream index segments from fast NVMe storage into RAM as needed, squeezing maximum performance out of affordable VPS instances.

This technical guide details the exact architecture, configuration parameters, and optimization strategies required to deploy a production-grade, large-scale vector search system using pgvector and HNSW on a constrained memory footprint.

---

1. The Architecture of HNSW in pgvector: Why RAM Exhaustion Occurs

To optimize the system, we must first understand how HNSW operates inside PostgreSQL. HNSW constructs a multi-layered graph structure where the bottom layer contains all vectors, and upper layers act as sparse skip-lists to accelerate search routing.

When executing a similarity query, pgvector traverses these graphs via random memory accesses. If the HNSW index exceeds the allocated PostgreSQL shared_buffers, the operating system is forced to fetch index pages from the disk. In a default configuration, random disk I/O introduces severe latency spikes, causing query performance to degrade from milliseconds to seconds.

Key Challenge: For 1,000,000 vectors of 1536 dimensions using float4 data types, the raw vector data alone requires approximately 6 GB. When you build an HNSW index with standard parameters (e.g., $M=16$, $EfConstruction=64$), the index structure can easily double or triple that footprint. Adding PostgreSQL overhead means a 16GB RAM VPS will rapidly trigger Out-Of-Memory (OOM) errors.
---

2. The Solution: Memory-Mapped Files (mmap) and the OS Buffer Cache

Instead of scaling vertically to expensive high-RAM bare metal servers, we can exploit how Linux manages file systems via Memory-Mapped Files (mmap). PostgreSQL inherently relies on the operating system's page cache to manage data blocks that do not fit into its internal shared_buffers.

When configured correctly, the Linux kernel treats the pgvector HNSW index files on disk as an extension of physical memory. When a specific node of the HNSW graph is requested during a query:

  • The kernel checks if the specific page is in the OS buffer cache.
  • If present (a cache hit), it reads directly from RAM at near-native speeds.
  • If absent (a cache miss), it triggers a minor page fault and streams the block from NVMe storage via demand paging.

To make this architecture viable on a VPS, the underlying storage must be high-performance NVMe SSDs. Because HNSW relies on random access, traditional HDDs or slow cloud block storage will completely kill performance.

---

3. Step-by-Step Configuration for Memory-Tiered pgvector

Implementing this optimization strategy requires a synchronized tuning approach across PostgreSQL internal parameters, index construction definitions, and Linux kernel settings.

Step 3.1: Optimizing PostgreSQL Memory Allocation

Counter-intuitively, when relying on memory-mapped files via the OS page cache for large indexes, you should not allocate all your VPS RAM to PostgreSQL's shared_buffers. If shared_buffers is too large, the OS has less room for the file cache, leading to duplication of data structures and inefficient memory recycling.

For an 8-core, 16GB RAM VPS handling a large vector dataset, apply the following adjustments in postgresql.conf:

# Allocate 25% of total RAM to shared_buffers
shared_buffers = 4GB

# Allow deep nested graph construction to use sufficient RAM during index builds
maintenance_work_mem = 4GB

# Utilize standard query workspace memory
work_mem = 64MB

# Tell the optimizer how much RAM is available via the OS cache
effective_cache_size = 11GB

# Optimize for random NVMe access patterns
random_page_cost = 1.1
effective_io_concurrency = 200

Step 3.2: Building Memory-Efficient HNSW Indexes

When executing the CREATE INDEX statement in pgvector, the parameters selected directly dictate the structural density of the HNSW graph and, consequently, its memory footprint. We can balance the recall-to-memory trade-off using the $M$ and $EfConstruction$ arguments.

Use the following SQL configuration to build an HNSW index optimized for tiered memory:

CREATE INDEX ON items USING hnsw (embedding vector_cosine_ops)
WITH (
  m = 16, 
  ef_construction = 64
);

Where:

  • m = 16: Limits the maximum number of bidirectional connection links per node. Lowering this from the default 32 reduces the index file size on disk by up to 40%, ensuring it fits better within the Linux page cache.
  • ef_construction = 64: Controls the entry evaluation bounds during index construction. A value of 64 maintains high index build speeds and avoids over-clustering structural paths that cause excessive random page faults later.
---

4. Advanced Vector Compression: Quantization

To further scale your VPS without upgrading hardware, you can couple memory-mapping with vector quantization. pgvector supports element-level dimensionality optimizations via expression indexes or external half-precision (fp16) casting if supported by your workflow.

By reducing the precision of the vector elements stored within the HNSW index from 32-bit floating points (float4) to 16-bit floating points, you instantly cut the memory footprint of your graph in half. This means twice as many vector nodes can reside within the same OS page cache allocation, drastically dropping the number of disk-bound page faults during high-concurrency vector searches.

---

5. Benchmarking and Monitoring Tiered Memory Performance

Once deployed, you must monitor whether your VPS is handling the memory-mapped index effectively or bottlenecking on disk I/O. Execute your target similarity search queries while monitoring the operating system using standard Linux utilities.

Monitoring Tools

  • vmstat: Execute vmstat 1 and monitor the bi (blocks in) and bo (blocks out) columns. High, continuous bi throughput during queries indicates frequent cache misses, meaning your vector index is heavily dependent on disk fetching.
  • pg_stat_database: Run queries against this internal PostgreSQL view to verify cache hit ratios. Aim for a local cache hit ratio above 90% across your active sessions.

Query Tuning

During runtime execution, you can dynamically tune the hnsw.ef_search variable. Increasing this value improves recall accuracy but causes the search algorithm to traverse more graph paths, increasing the likelihood of hitting unmapped disk pages:

SET hnsw.ef_search = 32; -- Fast performance, lower memory page requests
SELECT * FROM items ORDER BY embedding <=> '[0.012, -0.023, ...]' LIMIT 10;
---

Conclusion: High-Performance Vector Search Doesn't Require High Cost

Scaling large-scale vector search functionality using pgvector does not necessitate migrating to expensive, dedicated cloud vector databases or high-tier memory instances. By intelligently structuring your environment to use smaller $M$ parameters, shifting index caching dynamics onto the Linux kernel's memory-mapped capabilities, and running on fast NVMe-backed VPS nodes, you can achieve single-digit millisecond query latencies at a fraction of the traditional infrastructure cost.

As you continue to build out AI capabilities, keep this system blueprint in mind to maintain lean, efficient, and highly performant database architectures.

Scaling Large-Scale Vector Search on a Budget: Optimizing VPS with Memory-Mapped Files and HNSW in pgvector | DPTCloud