Self-Hosting a Graph-RAG AI Platform: Integrating Neo4j, pgvector, and Ollama on a Linux VPS
Introduction: The Evolution of Retrieval-Augmented Generation
In the rapidly evolving landscape of enterprise artificial intelligence, Retrieval-Augmented Generation (RAG) has emerged as the standard for grounding Large Language Models (LLMs) in proprietary data. Traditional RAG systems rely primarily on vector databases to perform semantic searches. While highly effective for finding isolated text chunks, standard vector RAG often falls short when dealing with complex, interconnected data structures or queries that require holistic, multi-hop reasoning.
This is where Graph-RAG comes into play. By combining the semantic capabilities of vector search with the structural intelligence of a knowledge graph, Graph-RAG enables AI systems to understand not just the words in your documents, but the intricate relationships between concepts, entities, and organizations. For enterprises looking to deploy this capability securely, self-hosting on a Linux Virtual Private Server (VPS) offers the ultimate balance of data privacy, infrastructure control, and cost predictability. In this comprehensive guide, we will walk through the architecture and deployment of a self-hosted Graph-RAG platform utilizing three industry-leading open-source technologies: Neo4j, pgvector, and Ollama.
---Understanding the Core Architecture
Before diving into the technical implementation, it is crucial to understand how these three components synergize to create a unified Graph-RAG pipeline:
- Ollama: Serves as our localized inference engine, running state-of-the-art open-source LLMs (such as Llama 3 or Mistral) and embedding models directly on your VPS, entirely eliminating third-party API dependencies.
- pgvector (PostgreSQL): Acts as our high-performance vector store, holding the dense embedding vectors generated by Ollama for rapid semantic similarity searches.
- Neo4j: Functions as our robust native graph database, mapping the explicit entities, properties, and relationships extracted from your data corpus.
By connecting these layers, your system can execute hybrid queries. It uses pgvector to locate relevant entry points based on semantic similarity, and then leverages Neo4j to traverse the surrounding relational graph, delivering an unparalleled depth of context to the LLM running on Ollama.
---Prerequisites and Environment Setup
To follow this guide, you will need a Linux VPS (Ubuntu 22.04 LTS or 24.04 LTS recommended). Given that we are running localized embeddings and LLMs, ensure your server meets the following minimum specifications:
- CPU: Modern 4-core or 8-core processor (with AVX2 support).
- RAM: Minimum 16 GB (32 GB highly recommended for fluid graph operations and LLM inference).
- Storage: 100 GB+ NVMe SSD.
- Docker & Docker Compose: Installed and configured on the system.
System Note: While a dedicated GPU (like an NVIDIA T4 or A10G) drastically accelerates LLM token generation, this setup can run entirely on CPU using Ollama's optimized CPU inference backends, making it highly accessible for testing and mid-scale deployments.---
Step-by-Step Deployment Guide
Step 1: Deploying the Containerized Infrastructure
The most efficient and isolated method to manage our stack is via Docker Compose. Create a new directory and define your infrastructure in a docker-compose.yml file:
version: '3.8'
services:
postgres:
image: pgvector/pgvector:pg16
container_name: postgres_vector
environment:
POSTGRES_USER: rag_user
POSTGRES_PASSWORD: SecurePassword123
POSTGRES_DB: rag_vector_db
ports:
- "5432:5432"
volumes:
- pgdata:/var/lib/postgresql/data
networks:
- rag_network
neo4j:
image: neo4j:5.19-community
container_name: neo4j_graph
ports:
- "7474:7474"
- "7687:7687"
environment:
- NEO4J_AUTH=neo4j/SecurePassword123
- NEO4J_PLUGINS=["apoc"]
volumes:
- neo4j_data:/data
- neo4j_logs:/logs
networks:
- rag_network
ollama:
image: ollama/ollama:latest
container_name: ollama_engine
ports:
- "11434:11434"
volumes:
- ollama_storage:/root/.ollama
networks:
- rag_network
volumes:
pgdata:
neo4j_data:
neo4j_logs:
ollama_storage:
networks:
rag_network:
driver: bridgeRun docker compose up -d to initialize all services in the background.
Step 2: Preparing Ollama Models
Once the containers are running, download the necessary models from the Ollama library. We need a model for text embedding and a model for generating the final response. Execute the following commands on your server host:
# Download the LLM for generation
docker exec -it ollama_engine ollama run llama3
# Download the embedding model
docker exec -it ollama_engine ollama pull nomic-embed-textStep 3: Initializing the Vector Database Extension
Connect to your PostgreSQL container to activate the pgvector extension and create a table tailored for storing chunked text data alongside their high-dimensional vector representations:
docker exec -it postgres_vector psql -U rag_user -d rag_vector_db -c "CREATE EXTENSION IF NOT EXISTS vector;"With the extension loaded, you can now construct your schemas, defining columns with the vector(768) datatype to perfectly match the 768 dimensions output by the nomic-embed-text model.
Connecting the Components with Python
To orchestrate the data ingestion pipeline, we use a Python script. This script ingests unstructured documents, chunks the text, creates relationships for Neo4j, generates embeddings via Ollama, and indexes those embeddings inside pgvector.
The Ingestion Process
First, ensure you have the required libraries installed in your Python environment: pip install neo4j psycopg2-binary requests langchain-community.
The orchestration workflow follows this programmatic sequence:
- Document Chunking: Raw text is divided into manageable paragraphs or semantic blocks.
- Entity & Relationship Extraction: Utilizing the LLM or a specialized NLP library, key entities (e.g., "Company A", "Product B") and their interactions ("Company A manufactures Product B") are identified.
- Graph Insertion: These entities are inserted into Neo4j using Cypher queries, forming the foundational structure of the knowledge graph.
- Vector Generation & Storage: The raw text chunk is passed to Ollama's embedding API, and the resulting vector is stored in pgvector, with a foreign key referencing the corresponding Neo4j node IDs.
Executing a Hybrid Graph-RAG Query
The true power of this architecture shines during retrieval. When a user submits a complex business query, the system executes a dual-pathway retrieval strategy:
Path A (Semantic Search): The query is converted into a vector via Ollama and submitted to pgvector. The database returns the top 3 most semantically relevant text chunks using cosine distance.
Path B (Structural Search): Utilizing the entities discovered in Path A, the system queries Neo4j to pull the adjacent graph topology—fetching related nodes, historical definitions, and multi-step linkages that standard text chunks miss.
The combined context (vector text chunks + graph metadata strings) is wrapped into a comprehensive prompt and sent to Ollama's llama3 model. The resulting output is not just statistically relevant, but contextually and relationally flawless.
Security, Optimization, and Production Considerations
Moving a self-hosted Graph-RAG setup from proof-of-concept to production requires careful attention to system design and maintenance:
- Memory Management: PostgreSQL, Neo4j, and Ollama are all memory-intensive. Allocate specific memory limits on your Docker containers to prevent the Linux kernel from triggering the Out-Of-Memory (OOM) killer.
- Index Optimization: For large datasets in pgvector, transition from flat index searching to an HNSW (Hierarchical Navigable Small World) index to maintain sub-millisecond retrieval speeds.
- Network Security: Do not expose ports
5432,7474, or11434to the public internet. Use an SSH tunnel, a strict UFW firewall configuration, or a secure VPN mesh network like Tailscale to communicate with your VPS services securely.
Conclusion
By self-hosting a Graph-RAG platform with Neo4j, pgvector, and Ollama on a Linux VPS, you break free from restrictive commercial licensing fees and data-sharing privacy risks. This architecture gives your business the cognitive power of localized LLMs alongside the structural precision of graph networks, providing a robust, highly reliable, and entirely private AI infrastructure tailored to your corporate intelligence needs.
