Building a Self-Hosted Personal AI Search Engine: Perplexica and SearXNG on a Budget 2GB RAM VPS
Introduction: The Rise of AI Search and the Privacy Dilemma
The landscape of information retrieval is undergoing a massive paradigm shift. Traditional search engines that present pages of links are rapidly being superseded by AI-powered answer engines like Perplexity AI. These platforms don't just find information; they synthesize, contextualize, and cite sources to deliver direct answers. However, relying on proprietary cloud services introduces significant concerns regarding data privacy, data logging, and subscription costs.
For professionals, researchers, and enterprises handling sensitive data, the ideal solution is a self-hosted alternative. But can you run a modern, LLM-powered AI search engine on a budget-friendly Virtual Private Server (VPS) with only 2GB of RAM? The answer is yes. By combining Perplexica (an open-source AI search engine) with SearXNG (a privacy-respecting metasearch engine) and leveraging external API-based Large Language Models (LLMs), you can deploy a robust, private search ecosystem for pennies on the dollar. This guide provides a step-by-step blueprint to achieve exactly that.
---Understanding the Architecture: Perplexica + SearXNG
Before diving into the technical installation, it is crucial to understand how these two components interact and why this specific combination is so powerful for resource-constrained environments like a 2GB RAM VPS.
- Perplexica: This acts as the frontend intelligence and orchestration layer. It takes the user's query, determines the search intent, interacts with the metasearch engine, parses the results, and passes them to an LLM to generate a coherent, cited response.
- SearXNG: This serves as the data aggregation engine. Instead of scraping Google or Bing directly (which triggers CAPTCHAs), SearXNG acts as a proxy metasearch engine. It aggregates results from over 70 search engines simultaneously, strips out tracking cookies, and returns clean JSON data to Perplexica.
Why this works on 2GB RAM: Running an LLM locally (like a 7B or 8B parameter model via Ollama) requires at least 8GB to 16GB of VRAM/RAM. By configuring Perplexica to use external, low-cost API endpoints (such as OpenAI, Groq, Anthropic, or OpenRouter), the local VPS only needs to handle network I/O, UI rendering, and basic data parsing. This keeps the memory footprint well under the 2GB threshold.
---Prerequisites and Environment Setup
To successfully follow this guide, ensure you have the following resources ready:
- A VPS running Ubuntu 22.04 LTS or 24.04 LTS with 1 vCPU, 2GB RAM, and at least 20GB of SSD storage.
- A domain or subdomain pointed to your VPS IP address (optional, but highly recommended for SSL setup via Nginx/Caddy).
- An API key from a fast LLM provider. We highly recommend Groq (for near-instantaneous speeds with Llama 3) or OpenRouter.
Step 1: System Update and Core Dependencies
Connect to your VPS via SSH and execute the following commands to update the system packages and install necessary utilities:
sudo apt update && sudo apt upgrade -y
sudo apt install curl git build-essential ufw -yStep 2: Install Docker and Docker Compose
Since both Perplexica and SearXNG are distributed as containerized applications, Docker is required. Install the latest Docker engine via the official convenience script:
curl -fsSL [https://get.docker.com](https://get.docker.com) -o get-docker.sh
sudo sh get-docker.shVerify the installation by checking the versions:
docker --version
docker compose version---Step-by-Step Deployment Guide
We will deploy both applications using Docker Compose to ensure clean networking and isolation.
1. Configuring SearXNG
First, create a dedicated directory for your search stack and set up the SearXNG configuration:
mkdir -p ~/ai-search/searxng && cd ~/ai-search/searxngCreate a settings.yml file to configure SearXNG to output JSON formatted data, which Perplexica requires:
nano settings.ymlPaste the following minimal, optimized configuration:
search:formats:
- html
- json
server:
port: 8080
bind_address: "0.0.0.0"
secret_key: "generate_a_secure_random_string_here"
ui:
theme: simple
engines:
- name: google
engine: google
shortcut: google
- name: bing
engine: bing
shortcut: bing
Make sure to replace the secret_key with a unique random hex string. You can generate one using the command: openssl rand -hex 32.
2. Cloning and Configuring Perplexica
Navigate back to your root deployment folder and clone the official Perplexica repository:
cd ~/ai-search
git clone [https://github.com/ItzCrazyK0/Perplexica.git](https://github.com/ItzCrazyK0/Perplexica.git)
cd PerplexicaPerplexica relies on a sample environment file. Copy it to create your active configuration:
cp sample.config.toml config.tomlEdit the config.toml file to hook into your external LLM provider and point to your local SearXNG instance:
nano config.tomlModify the following key-value pairs within the configuration file:
[GENERAL]---
PORT = 3000
SIMILAR_TO_PERPLEXITY = true
[SEARCH_ENGINE]
SEARXNG_URL = "http://searxng:8080"
[LLM]
PROVIDER = "groq" # Or openai, openrouter
KEY = "your_groq_api_key_here"
MODEL = "llama3-70b-8192"
3. Orchestrating with Docker Compose
To tie both services together under a unified local network, create a master docker-compose.yaml file inside the ~/ai-search directory:
cd ~/ai-search
nano docker-compose.yamlInsert the following configuration structure:
version: '3.8'services:
searxng:
image: searxng/searxng:latest
container_name: searxng
volumes:
- ./searxng/settings.yml:/etc/searxng/settings.yml
ports:
- "8080:8080"
restart: unless-stopped
networks:
- search_network
perplexica:
image: itzcrazyk0/perplexica:latest
container_name: perplexica
volumes:
- ./Perplexica/config.toml:/app/config.toml
ports:
- "3000:3000"
depends_on:
- searxng
restart: unless-stopped
networks:
- search_network
networks:
search_network:
driver: bridge
4. Launching the Stack
With all configuration files safely in place, launch your personal AI Search Engine using the detached mode command:
docker compose up -dVerify that both containers are running optimally by checking their real-time execution logs:
docker compose ps---Optimizing for a 2GB RAM VPS Environment
Running containers on low-spec infrastructure requires active resource conservation. To ensure your system remains entirely stable and avoids triggering the Linux Out-Of-Memory (OOM) killer, apply these optimization techniques:
1. Configure a Swap File
A swap file acts as an emergency overflow valve when your physical RAM is completely exhausted. For a 2GB RAM server, adding a 2GB or 4GB swap space is highly advantageous:
sudo fallocate -l 4G /swapfile
sudo chmod 600 /swapfile
sudo mkswap /swapfile
sudo swapon /swapfile
echo '/swapfile none swap sw 0 0' | sudo tee -a /etc/fstab2. Docker Log Rotation
Unchecked Docker container logs can grow exponentially over time, rapidly filling up small SSD capacities. Prevent this by appending log limits directly within your docker-compose.yaml file for each service profile:
logging:---
driver: "json-file"
options:
max-size: "10m"
max-file: "3"
Conclusion
By deploying Perplexica and SearXNG on a lean VPS, you successfully break free from proprietary data tracking paradigms while maintaining cutting-edge, context-aware AI search capabilities. This architecture proves that you do not need enterprise-grade, ultra-expensive local hardware rigs to enjoy complete digital sovereignty. For just a few dollars a month in basic cloud hosting fees, you now possess a fully customized, securely sandboxed, and extraordinarily fast AI answer engine tailored exclusively to your personal workflow requirements.
