Back to articles
Technology Insight

Building a Private Perplexity Alternative: Deploying DeepSeek-R1 and SearXNG on a 4GB RAM VPS

June 3, 2026

Introduction: The Quest for Private, AI-Powered Search

In the modern enterprise landscape, AI-powered search engines like Perplexity AI have revolutionized how we market research, analyze competitors, and synthesize complex information. However, relying on public cloud-based AI search engines poses a significant challenge for businesses: data privacy. Entering proprietary source code, confidential financials, or strategic queries into external platforms risks data leaks and regulatory non-compliance.

To mitigate these risks, organizations are increasingly turning to self-hosted alternatives. This technical guide demonstrates how to build your own "Private Perplexity" by orchestrating DeepSeek-R1 (a state-of-the-art reasoning model) and SearXNG (a privacy-respecting metasearch engine) on a highly cost-effective 4GB RAM Virtual Private Server (VPS). By leveraging Ollama, Docker, and specialized UI wrappers, you can achieve enterprise-grade, localized AI search infrastructure at a fraction of hyperscaler costs.

Architecture Overview: Maximizing Efficiency on Constrained Hardware

Running a localized reasoning LLM alongside a search aggregator on a 4GB RAM footprint requires extreme architectural optimization. Standard deployments of 7B or 14B parameter models will instantly crash due to Out-Of-Memory (OOM) errors. Our solution optimizes resources by integrating the following components:

  • DeepSeek-R1 (1.5B or 8B Distilled Q4 Quantized): We utilize the heavily optimized GGUF format via Ollama. The 1.5B model fits comfortably within 2GB of RAM, while the 8B model can be run using aggressive RAM swapping and aggressive quantization layers, though 1.5B offers the lowest latency on 4GB hardware.
  • SearXNG: A decentralized metasearch engine that aggregates results from over 70 search engines (Google, Bing, DuckDuckGo) without tracking user identities or storing search history.
  • Open WebUI or Perplexica: A responsive, user-friendly frontend interface that handles orchestration, issuing API queries simultaneously to SearXNG for web context and DeepSeek-R1 for reasoning.

By hosting this ecosystem inside an isolated Docker network on your private VPS, no queries or contextual documents ever leave your perimeter, ensuring total sovereignty over your corporate intelligence assets.

Prerequisites and Environment Setup

Before initiating the installation, ensure your infrastructure meets the following baseline requirements:

  • VPS Specifications: 2 vCPUs (Intel/AMD), 4GB RAM, and at least 40GB of SSD storage running Ubuntu 22.04 LTS or 24.04 LTS.
  • Swap Space: A minimum of 4GB to 8GB of configured virtual memory swap to absorb peak allocation spikes during concurrent model execution and web scraping.
  • Network: Inbound ports 80/443 (for the web UI) open, with a registered domain name pointing to your VPS IP for SSL configuration.

Step 1: System Optimization and Swap Allocation

First, log into your VPS via SSH and initialize system updates while allocating swap space to ensure memory stability:

sudo apt update && sudo apt upgrade -y
sudo fallocate -l 6G /swapfile
sudo chmod 600 /swapfile
sudo mkswap /swapfile
sudo swapon /swapfile
echo '/swapfile none swap sw 0 0' | sudo tee -a /etc/fstab
sudo sysctl vm.swappiness=10

Setting vm.swappiness=10 instructs the Linux kernel to prioritize physical RAM allocations, utilizing the swap file only to prevent immediate system OOM crashes.

Deploying the Search Core: SearXNG Configuration

We deploy SearXNG via Docker Compose. It acts as our privacy-preserving middleware, scraping web indices and formatting the output into clean JSON vectors that our local LLM can parse.

Step 2: Install Docker and Define Services

Execute the official Docker installation script and set up your project workspace:

curl -fsSL [https://get.docker.com](https://get.docker.com) -o get-docker.sh
sudo sh get-docker.sh
mkdir -p ~/private-perplexity && cd ~/private-perplexity

Create a docker-compose.yml file to orchestrate both SearXNG and the database cache backend (Redis) required for high-frequency search responses:

version: '3.8'

services:
  redis:
    image: redis:7-alpine
    command: redis-server --save "" --appendonly no
    networks:
      - perplexity-net
    restart: always

  searxng:
    image: searxng/searxng:latest
    volumes:
      - ./searxng:/etc/searxng:rw
    ports:
      - "8080:8080"
    networks:
      - perplexity-net
    depends_on:
      - redis
    restart: always

networks:
  perplexity-net:
    driver: bridge

Ensure you generate an automated secret key inside your ./searxng/settings.yml configuration file to authenticate queries securely before executing docker compose up -d.

Integrating DeepSeek-R1 via Ollama

Ollama allows us to serve quantized LLMs with minimized CPU/GPU overhead. Since our budget VPS lacks a dedicated GPU, Ollama will execute DeepSeek-R1 purely over highly optimized CPU instruction sets (AVX2).

Step 3: Download Ollama and Pull the Model

Install Ollama directly to the host system to minimize Docker-to-host resource translation overhead:

curl -fsSL [https://ollama.com/install.sh](https://ollama.com/install.sh) | sh

Once installed, pull the specialized DeepSeek-R1 distillation model. For a 4GB RAM threshold, the 1.5-billion parameter model is highly recommended due to rapid token processing speeds:

ollama run deepseek-r1:1.5b

This model integrates distinct tokens, allowing the local AI to structurally map out logical steps before cross-referencing live data streams provided by SearXNG.

Wiring the Frontend: The Perplexity Interface

To recreate the exact search experience of Perplexity—where a query triggers web searches, summarizes links, and lists foundational citations—we deploy Perplexica or Open WebUI with Web-Search enabled.

For this architecture, we utilize Open WebUI due to its native integrations with Ollama and robust multi-user authentication controls, essential for internal business teams.

Step 4: Launching Open WebUI with Web Access

Append the Open WebUI container definition into your existing docker-compose.yml file:

  open-webui:
    image: ghcr.io/open-webui/open-webui:main
    ports:
      - "3000:8080"
    volumes:
      - open-webui:/app/backend/data
    environment:
      - OLLAMA_BASE_URL=[http://host.docker.internal:11434](http://host.docker.internal:11434)
      - ENABLE_RAG=True
      - SEARXNG_QUERY_URL=http://searxng:8080/search?q=
extra_hosts: - "host.docker.internal:host-gateway" networks: - perplexity-net restart: always volumes: open-webui:

Run docker compose up -d to launch the complete system. Navigate to http://your-vps-ip:3000, configure an admin account, and enable the global web search toggle under the administration panel.

Performance Optimization and Enterprise Considerations

Running an operational AI framework on restricted infrastructure requires continuous maintenance. Implement these operational practices to maintain high availability:

  • Query Rate Limiting: Use Nginx or Caddy reverse proxies to throttle external requests, ensuring that multiple concurrent enterprise users do not lock up the CPU threads during the DeepSeek inference cycle.
  • Context Window Management: Set the maximum context window size in the Open WebUI interface to 4096 tokens. Expanding beyond this threshold exponentially increases RAM utilization, risking system degradation.
  • Automated Cache Purging: Set a cron job to flush old Redis database records weekly to maintain disk storage health and maximize system read/write IOPS.

Conclusion: Full Data Autonomy on a Budget

Building an autonomous, private alternative to Perplexity AI is no longer reserved for large enterprises with massive capital budgets. By orchestrating DeepSeek-R1 and SearXNG on a budget 4GB RAM VPS, your organization gains a highly secure, private intelligence asset. Your competitive analysis, sensitive IP checks, and market research remain completely confidential, safely hidden within an infrastructure you fully own and control. Start reclaiming your data privacy today by deploying your own private AI search mesh.