Self-Hosting a Secure Enterprise Document Management and RAG Search Engine via Onyx and pgvector on Docker VPS
Introduction: The Enterprise Knowledge Challenge
In the modern corporate landscape, data is an organization's most valuable asset. However, as an enterprise grows, its knowledge base inevitably fragments across scattered PDFs, internal wikis, spreadsheets, and isolated cloud drives. Employees waste countless hours searching for specific contract clauses, technical specifications, or compliance guidelines.
While generative AI and Large Language Models (LLMs) offer a transformative solution through Retrieval-Augmented Generation (RAG), sending proprietary corporate intelligence to public cloud APIs poses severe data privacy, compliance, and financial risks. The solution? Self-hosting.
This comprehensive guide will walk you through deploying a production-ready, self-hosted enterprise document management and RAG search platform using Onyx (an enterprise-grade knowledge platform) and pgvector (PostgreSQL's vector extension), entirely containerized via Docker on a Virtual Private Server (VPS).
Why This Stack? Architectural Advantages for Business
Choosing the right open-source stack is critical for performance, scalability, and ease of maintenance. Here is why this combination provides an institutional-grade architecture:
- Onyx (Enterprise Knowledge Interface): Provides a polished, business-ready UI/UX for user management, document ingestion, connectors (Slack, Google Drive, Confluence), and fine-grained access control. It abstracts complex RAG pipelines into a manageable control plane.
- pgvector (PostgreSQL Vector Database): Instead of adding another complex database to your infrastructure, pgvector allows you to store both traditional relational metadata and high-dimensional vector embeddings in a single, robust PostgreSQL database. It scales efficiently and ensures ACID compliance.
- Docker & VPS: Containerization guarantees environmental consistency, fast rollbacks, and effortless migration. Running on a private VPS ensures absolute data sovereignty, keeping your intellectual property completely within your perimeter.
Prerequisites and Infrastructure Sizing
Before launching the deployment, ensure your VPS meets the minimum hardware requirements. Because embedding generation and vector search can be CPU-intensive, we recommend the following baseline configuration:
Recommended VPS Specs: Minimum 4 vCPUs, 8GB RAM (16GB recommended for heavy ingestion), 100GB+ NVMe SSD storage, and Ubuntu 24.04 LTS or newer.
Additionally, ensure you have Docker and Docker Compose installed on your server, along with a domain name pointing to your VPS IP address for secure SSL configuration.
Step-by-Step Deployment Guide
Let us break down the installation process into actionable phases, from configuring the database to initializing the web interface.
Step 1: Setting Up the Docker Directory Structure
Connect to your VPS via SSH and create a structured directory to manage your persistent volumes and configuration files:
mkdir -p enterprise-rag/{postgres_data,onyx_data,nginx_config}
cd enterprise-rag
Step 2: Configuring the docker-compose.yml File
Create a docker-compose.yml file in your directory. This orchestration file configures PostgreSQL with the pgvector extension and pairs it with the Onyx backend and frontend services. Below is a highly optimized configuration blueprint:
Note: Ensure you replace placeholder passwords with strong, unique credentials before deploying.
version: '3.8'
services:
db:
image: pgvector/pgvector:pg16
container_name: enterprise_vector_db
restart: always
environment:
POSTGRES_USER: onyx_admin
POSTGRES_PASSWORD: StrongEnterprisePassword2026
POSTGRES_DB: onyx_knowledge
volumes:
- ./postgres_data:/var/lib/postgresql/data
ports:
- "5432:5432"
onyx-backend:
image: onyxai/onyx-backend:latest
container_name: onyx_backend
restart: always
depends_on:
- db
environment:
- DOCUMENT_INDEX_NAME=enterprise_index
- VECTOR_DB_TYPE=pgvector
- PGVECTOR_HOST=db
- PGVECTOR_PORT=5432
- PGVECTOR_USER=onyx_admin
- PGVECTOR_PASSWORD=StrongEnterprisePassword2026
- PGVECTOR_DB=onyx_knowledge
- AUTH_TYPE=basic
- WEB_DOMAIN=[https://knowledge.yourcompany.com](https://knowledge.yourcompany.com)
volumes:
- ./onyx_data:/app/storage
onyx-frontend:
image: onyxai/onyx-frontend:latest
container_name: onyx_frontend
restart: always
depends_on:
- onyx-backend
ports:
- "80:3000"
Step 3: Launching the Services
With the configuration file established, initialize the containers in detached mode by executing the following command:
docker compose up -d
Verify that all containers are operating optimally by reviewing the real-time process logs: docker compose ps and docker compose logs -f.
Configuring Your RAG Pipeline for Production
Once the platform is running, navigate to your server's IP address or configured domain via a web browser to complete the administrator setup wizard.
1. Connecting the LLM and Embedding Models
To power the RAG pipeline, you must link an LLM provider and an embedding model. For an entirely self-hosted, localized setup, you can point Onyx to an internal Ollama instance running open-weights models like Llama-3 or Mistral. Alternatively, for superior reasoning capabilities, you can safely connect corporate accounts via secure API endpoints like Azure OpenAI or Anthropic clusters.
2. Knowledge Ingestion & Chunking Strategy
Upload your enterprise assets—such as operational manuals, compliance documents, or code documentation—via the interface. Onyx automatically utilizes pgvector to execute the underlying data lifecycle:
- Document Parsing: Extracting raw text from disparate file formats (.pdf, .docx, .md).
- Semantic Chunking: Fragmenting long documents into optimized, readable paragraphs to preserve contextual relevance.
- Vectorization: Converting text blocks into multi-dimensional vectors using your selected embedding model, then saving them securely into the
pgvectorinstance.
Security and Optimization Best Practices
Deploying a production platform requires strict adherence to security protocols. Consider implementing these critical safeguards:
- Enforce SSL/TLS Encryption: Never pass enterprise data over unencrypted HTTP connections. Deploy a reverse proxy like Nginx or Traefik alongside Let's Encrypt to mandate HTTPS.
- Implement Role-Based Access Control (RBAC): Ensure that sensitive documents (e.g., HR files, financial audits) are only queryable by authorized user groups within the Onyx permission dashboard.
- Database Indexing for Speed: As your document library scales into hundreds of thousands of segments, connect directly to PostgreSQL and construct an HNSW (Hierarchical Navigable Small World) index on your vector columns to drastically accelerate search query return speeds.
Conclusion: True Data Sovereignty Realized
By self-hosting your document management and RAG search platform with Onyx and pgvector on Docker, your enterprise achieves the perfect balance of modern AI efficiency and strict data confidentiality. You successfully eliminate recurring per-user SaaS fees, bypass risk-prone public API limitations, and empower your team with an accurate, lightning-fast internal search engine.
Take complete control of your corporate intelligence today. Deploying a private RAG infrastructure is no longer an elite privilege reserved for tech giants—it is an accessible, strategic asset for every forward-thinking business.
