Back to articles
Technology Insight

Building a Lightweight RAG System on a 2GB RAM VPS: A Guide Using LiteDB, Qwen2.5-Coder, and Ollama

June 4, 2026

Introduction to Lightweight RAG Architecture

In the rapidly evolving landscape of Artificial Intelligence, Retrieval-Augmented Generation (RAG) has emerged as a cornerstone architecture for organizations seeking to ground Large Language Models (LLMs) in proprietary data. However, a common misconception persists: that implementing a robust RAG system necessitates enterprise-grade infrastructure, expensive vector databases, and massive computational overhead.

This comprehensive guide dismantles that assumption. We will walk through the architecture and implementation of a highly efficient, production-ready RAG system tailored for resource-constrained environments, specifically a Virtual Private Server (VPS) with just 2GB of RAM. By strategically combining LiteDB (a serverless .NET document database), Ollama (a lightweight LLM runner), and the highly optimized Qwen2.5-Coder model, we can achieve impressive performance without compromising on accuracy or breaking the budget.

The Core Components: Why This Stack?

To operate within a strict 2GB RAM boundary, every component in our pipeline must be carefully selected for its low memory footprint and high operational efficiency. Let us break down the core technology stack:

  • Ollama: Ollama serves as our local model execution engine. It manages model weights efficiently, handles quantization, and exposes a clean REST API, allowing us to run advanced models locally with minimal background memory usage.
  • Qwen2.5-Coder (0.5B or 1.5B Parameter Variant): The Qwen2.5-Coder series from Alibaba represents a breakthrough in small language models (SLMs). Despite its compact size, it demonstrates exceptional reasoning, syntax comprehension, and contextual awareness, making it perfect for processing enterprise documentation while adhering to strict RAM constraints.
  • LiteDB: Instead of deploying heavy, memory-hungry vector databases like qdrant or Milvus which typically demand significant baseline RAM, we utilize LiteDB. As a serverless, single-data-file embedded document store for .NET, it offers lightning-fast indexing, zero daemon overhead, and native support for storing text chunks alongside their corresponding vector embeddings.

Step-by-Step Environment Setup

Before writing the application logic, we must prepare our 2GB RAM VPS environment. Operating in a constrained memory space requires proactive optimization of the operating system and background processes.

1. Optimizing Swap Space

On a 2GB RAM machine, configuring virtual memory is critical to prevent the Linux Out-Of-Memory (OOM) killer from terminating your application processes. We recommend allocating a 4GB swap file:

sudo fallocate -l 4G /swapfile
sudo chmod 600 /swapfile
sudo mkswap /swapfile
sudo swapon /swapfile

To ensure this configuration persists across system reboots, append the following line to your /etc/fstab file:

/swapfile swap swap defaults 0 0

2. Installing and Configuring Ollama

Install Ollama via the official automated script, which handles systemd service configuration automatically:

curl -fsSL [https://ollama.com/install.sh](https://ollama.com/install.sh) | sh

Once installed, pull the optimized Qwen2.5-Coder model. For a 2GB RAM environment, the 0.5B parameter variant is highly recommended as it leaves ample breathing room for the OS and database caching, though the 1.5B variant can be used if tightly optimized:

ollama pull qwen2.5-coder:0.5b

Additionally, we need an embedding model to convert our document text chunks into vector representations. The nomic-embed-text model is exceptionally accurate and lightweight:

ollama pull nomic-embed-text

Implementing the RAG Pipeline

With our infrastructure ready, we can now build the RAG application logic. The architecture follows a standard two-phase pipeline: Data Ingestion and Query Generation.

Phase 1: Document Ingestion and Vector Storage

During the ingestion phase, raw text documents are parsed, broken down into manageable segments, vectorized, and saved to our persistent storage. Here is the step-by-step breakdown:

  1. Chunking: Large documents are split into smaller fragments (e.g., 500 characters with a 50-character overlap) to ensure that the retrieved context remains highly specific and fits comfortably within the model's prompt window.
  2. Embedding Generation: Each text chunk is sent to the Ollama embedding API endpoint (/api/embeddings) using the nomic-embed-text model, returning a dense vector array.
  3. LiteDB Persistence: We store a structured document containing the raw text, metadata (source, page number), and the vector array directly into LiteDB. Because LiteDB is embedded within our application process, data retrieval incurs virtually zero network latency.

Phase 2: Retrieval and Generation (The RAG Loop)

When a end-user submits a query, the system executes the following operational workflow:

First, the user's query is converted into a vector embedding using the same nomic-embed-text model. Second, a similarity search is performed against the LiteDB collection. Since LiteDB allows us to execute custom indexing routines, we can quickly calculate the Cosine Similarity or Euclidean Distance between the query vector and our stored document vectors.

Mathematically, the Cosine Similarity between a query vector $A$ and a document vector $B$ is defined as:

$$\text{Similarity} = \frac{A \cdot B}{\|A\| \|B\|}$$

The top-k highest-scoring text chunks are retrieved and injected into a structured system prompt. Finally, this combined context is sent to Qwen2.5-Coder to generate an accurate, contextualized answer.

Memory Optimization Strategies for 2GB VPS

Running this entire pipeline smoothly on minimal hardware requires strict adherence to resource management best practices. Implement these strategies to maintain system stability:

  • Model Concurrency Limitation: Configure Ollama to handle requests sequentially rather than concurrently by setting the environment variable OLLAMA_NUM_PARALLEL=1. This prevents memory spikes when multiple queries arrive simultaneously.
  • Garbage Collection Tuning: If developing the orchestration layer in .NET or Node.js, configure the runtime for aggressive garbage collection and low-memory footprints (e.g., using workstation GC mode rather than server GC mode in .NET).
  • Streaming Responses: Always enable token streaming (stream: true) when calling the Ollama generation endpoint. This reduces memory buffering on both the server and client sides, delivering a faster Time-to-First-Token (TTFT).

Conclusion

Building a Retrieval-Augmented Generation system does not require capital-intensive cloud infrastructure or massive GPU clusters. By combining the efficiency of LiteDB, the compact intelligence of Qwen2.5-Coder, and the seamless orchestration of Ollama, you can deploy a highly functional, secure, and private RAG system on a standard 2GB RAM VPS. This setup provides an ideal, cost-effective blueprint for startups, hobbyists, and enterprises looking to prototype AI solutions without high operational overhead.

Building a Lightweight RAG System on a 2GB RAM VPS: A Guide Using LiteDB, Qwen2.5-Coder, and Ollama | DPTCloud