Back to articles
Technology Insight

Building a Secure Private AI Knowledge Base with AnythingLLM and Docker on a VPS

May 30, 2026

Introduction: The Enterprise Imperative for Data Privacy in the AI Era

In the modern corporate landscape, Artificial Intelligence (AI) and Large Language Models (LLMs) have transitioned from speculative innovations to essential drivers of operational efficiency. Enterprises routinely leverage these technologies to automate customer service, streamline internal operations, and analyze vast repositories of corporate documentation. However, this rapid adoption introduces a critical vulnerability: data privacy and security compliance.

When organizations utilize public cloud-based AI services, they often inadvertently expose proprietary data, trade secrets, and sensitive customer information to third-party providers. For enterprise legal departments, financial institutions, and healthcare providers, this risk is unacceptable. The solution is the architecture of a Private AI Knowledge Base—a system where data ingestion, embedding, vector storage, and model inference occur entirely within an infrastructure controlled exclusively by the organization. This comprehensive technical guide provides a step-by-step framework to deploy a secure, scalable, and private AI knowledge base using AnythingLLM and Docker on a Virtual Private Server (VPS).

---

1. Why Choose AnythingLLM, Docker, and a VPS?

Building a private AI ecosystem requires selecting a software stack that balances robust security, architectural flexibility, and ease of maintenance. The combination of AnythingLLM, Docker, and a dedicated VPS represents the industry standard for mid-sized enterprise deployments due to several distinct operational advantages:

  • AnythingLLM: An all-in-one desktop and enterprise application that transforms raw documents into a structured, chat-driven knowledge base. It natively supports multi-user management, strict permission controls, and seamlessly integrates with various LLM providers and vector databases. It abstracts the complexity of Retrieval-Augmented Generation (RAG) pipelines into a refined, enterprise-grade user interface.
  • Docker Containerization: Deploying complex AI applications directly onto a host operating system often introduces dependency conflicts and configuration drift. Docker isolates AnythingLLM and its auxiliary services into lightweight, predictable containers, ensuring consistent performance across development, staging, and production environments.
  • Virtual Private Server (VPS): Utilizing a dedicated VPS offers full root access, predictable monthly operational expenditure, and guaranteed resource allocation. More importantly, it ensures your proprietary data remains isolated from public shared hosting environments, satisfying strict data sovereignty requirements.
---

2. System Prerequisites and Architecture Planning

Before initiating the technical deployment, ensure your target VPS environment meets or exceeds the following architectural requirements to handle concurrent users and continuous document embedding processes:

Hardware Recommendations

  • CPU: Minimum 4 vCPUs (Intel Xeon or AMD EPYC optimized instances are highly recommended for handling embedding calculations).
  • RAM: 8 GB minimum; 16 GB or higher is strongly preferred if hosting an on-premise open-source LLM (such as Llama 3 via Ollama) on the same machine.
  • Storage: 50 GB+ NVMe SSD storage to accommodate OS system overhead, docker images, localized vector databases, and enterprise document repositories.
  • Operating System: Ubuntu Server 22.04 LTS or 24.04 LTS for maximum stability and widespread container support.

Network and Security Prerequisites

  1. A fully qualified domain name (FQDN) mapped via DNS A-Record to your VPS public IP address (e.g., ai.yourcompany.com).
  2. A properly configured firewall (UFW or cloud provider security groups) restricting traffic to essential ports only: 22 (SSH), 80 (HTTP), and 443 (HTTPS).
---

3. Step-by-Step Deployment Guide

Step 3.1: Server Initialization and Docker Installation

Establish a secure SSH connection to your VPS and execute a comprehensive system update to ensure all core libraries are secure and current:

sudo apt update && sudo apt upgrade -y

Next, install the official Docker engine along with the Docker Compose plugin utilizing the standard repository setup:

# Install prerequisites
sudo apt install -y curl apt-transport-https ca-certificates software-properties-common

# Add Docker's official GPG key
curl -fsSL [https://download.docker.com/linux/ubuntu/gpg](https://download.docker.com/linux/ubuntu/gpg) | sudo gpg --dearmor -o /usr/share/keyrings/docker-archive-keyring.gpg

# Set up the stable repository
echo "deb [arch=$(dpkg --print-architecture) signed-by=/usr/share/keyrings/docker-archive-keyring.gpg] [https://download.docker.com/linux/ubuntu](https://download.docker.com/linux/ubuntu) $(lsb_release -cs) stable" | sudo tee /etc/list.d/docker.list > /dev/null

# Install Docker Engine
sudo apt update
sudo apt install -y docker-ce docker-ce-cli containerd.io docker-buildx-plugin docker-compose-plugin

# Ensure Docker starts on system boot
sudo systemctl enable docker
sudo systemctl start docker

Step 3.2: Configuration of AnythingLLM Storage and Environment

To ensure persistent storage of user profiles, uploaded documents, and vector embeddings across container restarts, create a dedicated directory structure on the host filesystem:

# Create storage directory
sudo mkdir -p /var/lib/anythingllm

# Configure an environment file for explicit application control
sudo touch /var/lib/anythingllm/.env

Step 3.3: Launching the AnythingLLM Container via Docker Compose

Using docker-compose simplifies management, multi-container scaling, and future software updates. Create a production-grade docker-compose.yml file within an administrative directory:

mkdir -p ~/private-ai && cd ~/private-ai
nano docker-compose.yml

Populate the file with the following standard service configuration:

version: '3.8'

services:
  anythingllm:
    image: mintplexlabs/anythingllm:latest
    container_name: anythingllm
    restart: always
    ports:
      - "3001:3001"
    volumes:
      - /var/lib/anythingllm:/app/storage
      - /var/lib/anythingllm/.env:/app/storage/.env
    environment:
      - STORAGE_DIR=/app/storage
    logging:
      driver: "json-file"
      options:
        max-size: "10m"
        max-file: "3"

Deploy the container in detached background mode by executing:

sudo docker compose up -d

Verify that the container is actively running and healthy by inspecting the real-time operational logs:

sudo docker ps
sudo docker logs -f anythingllm

---

4. Implementing Production Security: Reverse Proxy and SSL Encryption

Accessing your knowledge base directly over an unencrypted port (3001) exposes administrative credentials and corporate data to interception. To mitigate this security risk, implement Nginx as a reverse proxy coupled with Let's Encrypt TLS certificates.

Step 4.1: Install Nginx

sudo apt install -y nginx
sudo systemctl enable nginx

Step 4.2: Configure Nginx Server Block

Create a dedicated configuration file for your AI knowledge base subdomain:

sudo nano /etc/nginx/sites-available/anythingllm

Insert the following production configuration, ensuring you substitute your actual domain name:

server {
    listen 80;
    server_name ai.yourcompany.com;

    client_max_body_size 100M; # Accommodates large corporate PDF uploads

    location / {
        proxy_pass [http://127.0.0.1:3001](http://127.0.0.1:3001);
        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;
    
        # WebSockets support for real-time text streaming
        proxy_http_version 1.1;
        proxy_set_header Upgrade $http_upgrade;
        proxy_set_header Connection "Upgrade";
    }
}

Enable the site and restart the Nginx service:

sudo ln -s /etc/nginx/sites-available/anythingllm /etc/nginx/sites-enabled/
sudo nginx -t
sudo systemctl restart nginx

Step 4.3: Automate SSL Provisioning via Certbot

Execute the following commands to install Certbot and automatically acquire and configure a secure Let's Encrypt SSL certificate:

sudo apt install -y certbot python3-certbot-nginx
sudo certbot --nginx -d ai.yourcompany.com --non-interactive --agree-tos --email [email protected]
---

5. Fine-Tuning Your Private Knowledge Base Configuration

With network security established, navigate to [https://ai.yourcompany.com](https://ai.yourcompany.com) via a web browser to complete the initial administrative setup wizard.

Model Selection & Embedding Management

AnythingLLM provides total flexibility regarding the underlying computational engines:

  • Fully Sovereign Deployment: Connect AnythingLLM to a localized instance of Ollama or LocalAI running on the same VPS or a separate internal GPU cloud. Selecting models like Llama 3 or Mistral guarantees no data ever leaves your hardware border.
  • Hybrid Deployment: If high performance takes priority over total isolation, configure AnythingLLM to interface securely via encrypted APIs with enterprise instances of OpenAI, Anthropic, or Azure OpenAI.

Vector Database Configuration

AnythingLLM features a high-performance built-in vector database (LanceDB) that works out-of-the-box without further manual configuration. For enterprise-grade scaling encompassing millions of pages of text, the system easily bridges to dedicated external vector stores such as ChromaDB, Pinecone, or Weaviate.

---

Conclusion: Establishing Modern Data Governance

Deploying a Private AI Knowledge Base using AnythingLLM and Docker on a dedicated VPS bridges the gap between revolutionary AI capabilities and uncompromising corporate security. By self-hosting this infrastructure, your business retains total ownership of its data pipeline, achieves compliance with international data privacy frameworks, and empowers internal teams with an intelligent, highly localized search and synthesis tool.

As next steps, establish automated nightly backup routines for your host directory (/var/lib/anythingllm) and implement strict multi-factor authentication (MFA) within the user administration panel to ensure your proprietary intelligence layer remains permanently resilient and secure.

Building a Secure Private AI Knowledge Base with AnythingLLM and Docker on a VPS | DPTCloud