Back to articles
Technology Insight

Deploying RagFlow on a VPS: Enterprise RAG for Complex Documents, Multi-Format Tables, and Scanned PDFs

June 2, 2026

Introduction: The Enterprise RAG Dilemma

In the era of corporate intelligence, Retrieval-Augmented Generation (RAG) has emerged as the gold standard for leveraging internal knowledge bases. However, many organizations quickly hit a wall when transitioning from basic proof-of-concepts to production-grade deployment. Standard RAG pipelines frequently fail when encountering real-world enterprise documents: multi-page financial reports with dense nested tables, complex multi-column legal frameworks, and low-quality scanned PDFs from legacy archives.

Traditional text-splitting chunking strategies treat documents as linear strings of characters, inadvertently shredding table structures and separating critical headers from their numerical data. The result? Halucinations, inaccurate metrics, and unreliable AI insights. RagFlow changes this paradigm. By utilizing advanced deep document understanding, RagFlow reconstructs the visual layout of files before vectorization, ensuring unprecedented accuracy. In this comprehensive guide, we will explore how to deploy RagFlow on a Virtual Private Server (VPS) to create a secure, self-hosted, enterprise-grade RAG solution.

Why RagFlow for Complex Corporate Documentation?

Before diving into the technical deployment, it is essential to understand why RagFlow is uniquely suited for complex business environments compared to conventional frameworks:

  • Deep Document Parsing (OCR & Layout Analysis): RagFlow does not just read raw text. It identifies visual boundaries, separating headers, footers, footnotes, images, and embedded charts.
  • Table Structure Reconstruction: It recognizes rows, columns, and merged cells, converting visual tables into structured formats that Large Language Models (LLMs) can actually comprehend.
  • Robust Handling of Scanned PDFs: Equipped with built-in, high-efficiency Optical Character Recognition (CR), it extracts legible text even from faded, tilted, or low-resolution scanned documents.
  • Hybrid Search Capabilities: It combines keyword-based BM25 retrieval with dense vector embeddings to ensure both precision and semantic relevance are optimized.

By hosting RagFlow on your own VPS, your enterprise retains 100% data sovereignty, mitigating the privacy risks associated with uploading confidential IP, financial ledgers, or employee data to third-party SaaS platforms.

Prerequisites and VPS Sizing Guidelines

RagFlow relies on sophisticated machine learning models for layout parsing, OCR, and embedding generation. Therefore, selecting a VPS with adequate compute power is crucial for acceptable inference speeds.

Minimum Requirements (For Testing & Small Datasets)

  • CPU: 4 Cores (Intel Xeon or AMD EPYC equivalent)
  • RAM: 16 GB
  • Storage: 100 GB NVMe SSD
  • OS: Ubuntu 22.04 LTS

Recommended Requirements (For Production & High-Volume OCR)

  • CPU: 8 Cores or more
  • RAM: 32 GB
  • GPU: Dedicated NVIDIA GPU (e.g., T4, A10G, or L4) with 16GB+ VRAM (Highly recommended for accelerating parsing and local LLM inference)
  • Storage: 500 GB+ NVMe SSD
Note: While RagFlow can run completely on CPU-only VPS configurations by utilizing external LLM APIs (like OpenAI, Anthropic, or Cohere), its internal layout analysis models still benefit significantly from GPU acceleration.

Step-by-Step VPS Deployment Guide

We will utilize Docker Compose for a streamlined, reproducible setup. Ensure your VPS firewall allows traffic on ports 80, 443, and any custom ports you intend to use.

Step 1: System Update and Dependency Installation

First, log into your VPS via SSH and update the system packages. Next, install Docker and Docker Compose if they are not already present.

sudo apt update && sudo apt upgrade -y
sudo apt install curl git software-properties-common -y

# Install Docker
curl -fsSL [https://get.docker.com](https://get.docker.com) -o get-docker.sh
sudo sh get-docker.sh

# Manage Docker as a non-root user (Optional but recommended)
sudo usermod -aG docker $USER

Step 2: Clone the RagFlow Repository

Clone the official RagFlow repository from GitHub and navigate into the deployment directory:

git clone [https://github.com/infiniflow/ragflow.git](https://github.com/infiniflow/ragflow.git)
cd ragflow/docker

Step 3: Configure Environment Variables

RagFlow provides a template file for environment configurations. Copy the template and adjust the variables based on your server architecture:

cp .env.example .env

Open the .env file using your preferred text editor (e.g., nano .env). Pay close attention to the following configurations:

  • RAGFLOW_IMAGE: Ensure you are targeting the correct version tags. If you are running on a CPU-only server, verify you aren't pulling the heavy CUDA-dependent image variant unless required.
  • HTTP_PORT: Define the external port for the web interface (default is usually 80 or 8000).
  • Modify database passwords (MySQL, Elasticsearch/MinIO) from their defaults to secure your environment.

Step 4: Launching the Containers

Because RagFlow integrates multiple robust components—including Elasticsearch for keyword indexing, MinIO for object storage, Redis for caching, and MySQL for relational data—the initial pull may take several minutes depending on your network bandwidth. Launch the stack using the following command:

sudo docker compose up -d

To monitor the startup progress and ensure all backend services initialize without errors, use the log viewer:

sudo docker compose logs -f

Configuring the Engine for Complex Enterprise Documents

Once all containers are successfully running, access the RagFlow interface by navigating to your VPS IP address (e.g., http://your_vps_ip) via a web browser. Create your administrator account to access the primary dashboard.

Optimizing the Knowledge Base for Tables and Layouts

When creating a new Knowledge Base within RagFlow, you will be prompted to select a Parser Configuration. This selection determines how your documents are interpreted:

  1. The 'Table' or 'General' Template: For standard financial reports, do not use naive text chunking. Choose the general document template or the specific table parsing pipeline. This instructs the backend vision models to preserve column headers alongside cell values.
  2. OCR Engine Selection: If your files consist of scanned PDFs, ensure the OCR language packs match your documentation language (e.g., English, multi-language corporate bundles) to prevent garbled character outputs.
  3. Embedding Model Selection: Connect your system to an embedding model provider. For maximum security, you can point RagFlow to a local Ollama instance running on the same VPS, or utilize high-dimensional enterprise APIs like OpenAI's text-embedding-3-large.

Best Practices for Maintaining Enterprise RAG Security

Deploying on a VPS gives you total control over infrastructure architecture. To safeguard proprietary corporate data, implement these secondary security measures immediately:

  • Reverse Proxy and SSL: Do not leave the raw HTTP port exposed to the open web. Set up an Nginx or Caddy reverse proxy and provision an SSL certificate via Let's Encrypt to enforce encrypted HTTPS traffic.
  • IP Whitelisting: Restrict access to the RagFlow dashboard port at the VPS firewall level (using ufw or your cloud provider's security groups) so that only internal corporate VPN IPs can reach the application.
  • Automated Backups: Schedule automated cron jobs to backup the mapped Docker volumes, specifically the MinIO object store containing your source documents, and the underlying vector database indexes.

Conclusion

Standard RAG architectures often fail when confronted with the messy reality of enterprise documentation. By deploying RagFlow on a dedicated VPS, businesses bridge the gap between basic semantic search and deep, layout-aware document intelligence. Whether your workflows depend on parsing intricate financial matrices, interpreting complex multi-column legal briefs, or extracting value from legacy scanned PDFs, RagFlow maintains structural integrity where other parsers fail—all while keeping your critical data securely under your own institutional control.

Deploying RagFlow on a VPS: Enterprise RAG for Complex Documents, Multi-Format Tables, and Scanned PDFs | DPTCloud