Enterprise GraphRAG Deployment: Scaling Vector-Graph Hybrids with Neo4j and Ollama on Linux VPS
The Evolution of RAG: Why Vector Search is No Longer Enough
Retrieval-Augmented Generation (RAG) has rapidly become the standard architecture for grounding Large Language Models (LLMs) in proprietary enterprise data. However, as production systems scale, traditional vector-based RAG architectures are exposing critical limitations. While vector databases excel at finding localized semantic similarities, they fundamentally struggle with structured relationships, global document summarization, and multi-hop reasoning. If an executive asks, 'What are the systemic risks across all supply chain contracts signed in Q3?', standard vector search fails to connect the dots across disparate text chunks.
Enter GraphRAG. By combining the semantic understanding of vectors with the explicit relationship mapping of a Knowledge Graph (KG), GraphRAG allows organizations to capture both the hierarchical structure and the intricate webs of connection within their data. In this comprehensive technical guide, we will walk through deploying a fully self-hosted, enterprise-grade GraphRAG system utilizing Neo4j as our graph and vector engine, and Ollama for localized LLM inference, all hosted on a secure Linux Virtual Private Server (VPS).
The Core Architecture: Neo4j, Ollama, and Localized Intelligence
Building an enterprise-ready system requires balancing performance, security, and cost. Relying on external APIs for LLM inference and vector embeddings often introduces regulatory compliance challenges, unpredictable latency, and spiraling operational costs. Our open-source, self-hosted stack mitigates these risks perfectly:
- Neo4j: The industry-leading graph database that natively supports both graph traversals and vector index searches simultaneously, serving as our unified retrieval engine.
- Ollama: A robust, lightweight framework for running advanced LLMs (like Llama 3 or Mistral) and embedding models locally, ensuring total data sovereignty.
- Linux VPS: An optimized environment (Ubuntu 22.04 LTS or newer) providing dedicated compute, memory, and storage allocation without the overhead of managed cloud services.
Enterprise Note: To handle parsing, entity extraction, and simultaneous graph queries efficiently, your Linux VPS should ideally be provisioned with at least 8 vCPUs, 32GB of RAM, and fast NVMe storage. If real-time, high-throughput inference is required, prioritizing a VPS with GPU passthrough (e.g., NVIDIA A10G or T4) is highly recommended.
Step-by-Step Deployment Blueprint on a Linux VPS
1. Environment Preparation and System Optimization
Before installing the core software, we must update the host system and configure essential kernel parameters to ensure Neo4j can handle heavy transaction volumes and concurrent connections seamlessly.
sudo apt update && sudo apt upgrade -y
sudo apt install -y curl git apt-transport-https ca-certificates htopNext, adjust the maximum open file descriptors by editing /etc/security/limits.conf to prevent 'too many open files' errors during massive graph indexing operations:
neo4j soft nofile 40000
neo4j hard nofile 600002. Deploying Neo4j Enterprise/Community Edition via Docker
Using Docker ensures clean isolation of the database engine and simplifies backup routines. We will expose both the bolt protocol (7687) and the browser interface (7474), while mounting persistent volumes to protect our data.
docker run -d \
--name neo4j-graphrag \
-p 7474:7474 -p 7687:7687 \
-v /opt/neo4j/data:/data \
-v /opt/neo4j/logs:/logs \
-e NEO4J_AUTH=neo4j/YourSecurePassword123! \
-e NEO4J_PLUGINS='["apoc"]' \
--restart unless-stopped \
neo4j:latestThe inclusion of the APOC (Awesome Procedures on Cypher) plugin is mandatory, as it provides the advanced data transformation functions required during the entity extraction and graph construction phases.
3. Installing Ollama and Provisioning Local Models
Install Ollama directly onto the host to maximize hardware utilization. Run the official setup script:
curl -fsSL [https://ollama.com/install.sh](https://ollama.com/install.sh) | shOnce installed, pull an advanced general-purpose model for entity extraction and a dedicated embedding model for generating vector representations. For enterprise applications, Llama3 (8B) offers an exceptional balance of speed and reasoning capability, while nomic-embed-text provides highly accurate contextual embeddings:
ollama pull llama3:8b
ollama pull nomic-embed-textImplementing the GraphRAG Pipeline: Knowledge Graph Construction
With infrastructure running, the system must ingest unstructured text documents, extract relevant nodes (Entities) and edges (Relationships), and index them into Neo4j. The enterprise GraphRAG pipeline follows a strict, highly organized workflow:
- Document Chunking: Raw text is parsed into overlapping chunks (e.g., 500 tokens with a 100-token overlap) to maintain contextual continuity across boundaries.
- Vector Embedding: Each text chunk is sent to Ollama’s
nomic-embed-textmodel to generate a high-dimensional vector. These are stored directly in Neo4j using a vector index. - Entity-Relationship Extraction: The text chunks are analyzed by
llama3:8bvia structured prompting. The model extracts key entities (e.g., Organization, Person, Product) and defines their exact relationship (e.g., ACQUIRED, DEVELOPED, PARTNERED_WITH). - Graph Resolution and Merging: Cypher queries ingest the extracted entities, automatically deduplicating nodes and building a interconnected knowledge web anchored to the original text chunks.
- Network Hardening: Never expose ports 7474, 7687, or 11434 directly to the public internet. Implement an Nginx reverse proxy secured by Let's Encrypt SSL certificates, and configure strict firewall rules via
ufwto allow access only from authorized enterprise VPN subnets. - Cache Management and Index Tuning: Create explicit indexes on frequently searched entity properties (e.g.,
CREATE INDEX FOR (n:Entity) ON (n.name)) to ensure traversal latencies remain sub-millisecond even as the graph grows to millions of nodes. - Backup Strategies: Implement scheduled cron jobs to trigger hot backups of Neo4j data directories, compressing them and transferring them to offsite enterprise object storage solutions securely.
To accelerate this layout in production, developers typically utilize frameworks like LangChain or LlamaIndex, connecting to Neo4j via the official Python driver and routing LLM calls directly to the local Ollama API endpoint (http://localhost:11434).
The Hybrid Retrieval Process: Executing Multi-Hop Queries
The true power of GraphRAG shines during execution. When an end-user submits a complex query, the system executes a multi-stage hybrid search strategy:
Phase 1: Vector Search
The user query is converted into an embedding and a vector similarity search is performed across the text chunk nodes in Neo4j. This quickly identifies the top candidate documents containing highly relevant semantic matches.
Phase 2: Graph Traversal
Starting from those highly ranked text chunks, the system executes graph traversal queries (using Cypher) to fetch adjacent entity nodes up to 2 or 3 hops deep. This pulls in structural context that was never explicitly mentioned in the matching text chunk itself, but is inherently connected across the broader knowledge graph.
Phase 3: Context Compilation and LLM Synthesis
The retrieved vector text chunks and the structured graph relationships are merged into a comprehensive context window. This highly rich, factual prompt is passed to Ollama, instructing the local LLM to generate an exhaustive, verified answer completely free of hallucinations.
Enterprise Security, Optimization, and Maintenance
Operating GraphRAG at a corporate scale requires strict adherence to system maintenance protocols and security standards:
Conclusion
By moving beyond isolated vector databases and establishing a self-hosted GraphRAG architecture with Neo4j and Ollama on a Linux VPS, your enterprise gains a monumental advantage. You build an intelligent system capable of deep, contextual reasoning and cross-document analysis, while retaining absolute control over your data privacy and operational infrastructure. The future of enterprise AI is deeply connected, highly secure, and structurally grounded.
