Back to articles
Technology Insight

Building a Self-Hosted AI Search Engine: Deploying Perplexica and SearXNG on a Budget 2GB VPS

May 30, 2026

Introduction: The Shift Toward Private, AI-Powered Search

The search engine landscape is undergoing a massive paradigm shift. Traditional, ad-driven keyword search is rapidly losing ground to AI-powered answer engines that synthesize information from multiple sources to deliver direct, conversational responses. While platforms like Perplexity AI have popularized this approach, relying entirely on proprietary ecosystems raises valid concerns regarding data privacy, search tracking, and mounting subscription costs.

For enterprises and tech-savvy professionals, the solution lies in self-hosting. This guide provides a comprehensive blueprint to architect your own private 'AI Search Engine' using two powerful open-source tools: Perplexica (the AI search frontend) and SearXNG (the privacy-respecting metasearch engine). Crucially, we will demonstrate how to optimize and deploy this entire stack on a highly affordable, constrained environment: a VPS with just 2GB of RAM.

---

Understanding the Architecture: Perplexica and SearXNG

Before diving into the technical deployment, it is vital to understand how these two components interact to deliver an enterprise-grade search experience without exhausting system resources.

  • SearXNG: Actings as our privacy-preserving metasearch engine, SearXNG aggregates search results from dozens of major platforms (Google, Bing, DuckDuckGo) while stripping out tracking cookies, user identifiers, and intrusive ads. It serves as the clean data retrieval layer.
  • Perplexica: This is the intelligent orchestration layer. It takes your natural language query, utilizes SearXNG to fetch the most relevant live web pages, reads the contents of those pages, and passes them to a Large Language Model (LLM) to synthesize a coherent, cited answer.
By decoupling the search aggregation (SearXNG) from the cognitive synthesis (Perplexica) and offloading the heavy AI processing to external API endpoints, we can run this sophisticated setup seamlessly on a lightweight 2GB RAM server.
---

Prerequisites and Environment Setup

To successfully execute this deployment, ensure you have the following prerequisites ready:

  1. A Cloud VPS: Running Ubuntu 24.04 LTS or 22.04 LTS with at least 1 vCPU and 2GB of RAM. Providers like Hetzner, DigitalOcean, or Linode are ideal.
  2. An External LLM API Key: Since running an LLM locally (like Llama 3) requires a minimum of 8GB–16GB of VRAM/RAM, we will use lightweight external APIs. Options include OpenAI (GPT-4o-mini), Groq (Llama 3 70B), or Anthropic. Groq is highly recommended for this setup due to its extreme speed and generous free/low-cost tiers.
  3. Docker and Docker Compose: Installed on your target VPS.

Step 1: System Optimization for 2GB RAM

Operating on a 2GB RAM threshold requires careful resource management. Without virtual memory modification, Docker containers may trigger the Linux Out-Of-Memory (OOM) killer during build or peak utilization phases. We mitigate this by establishing a 4GB Swap file.

sudo fallocate -l 4G /swapfile
sudo chmod 600 /swapfile
sudo mkswap /swapfile
sudo swapon /swapfile
echo '/swapfile none swap sw 0 0' | sudo tee -a /etc/fstab

Verify the swap allocation by executing free -h. You should now see 2GB of physical memory supplemented by 4GB of virtual swap space, providing an essential safety buffer.

---

Step 2: Deploying SearXNG as the Data Provider

We will deploy SearXNG via Docker. First, create a dedicated directory structure for our project to maintain organization and modularity:

mkdir -p ~/ai-search/searxng && cd ~/ai-search/searxng

Next, we must configure SearXNG to return output format types compatible with Perplexica. Create a configuration file named settings.yml:

# settings.yml excerpt
search:
  formats:
    - html
    - json
server:
  secret_key: "change_this_to_a_secure_random_string"
  bind_address: "0.0.0.0"
  port: 8080
ui:
  theme: simple

Now, construct the docker-compose.yml file specifically for SearXNG within the same directory:

version: '3.7'
services:
  searxng:
    image: docker.io/searxng/searxng:latest
    container_name: searxng
    volumes:
      - ./settings.yml:/etc/searxng/settings.yml:ro
    ports:
      - "127.0.0.1:8080:8080"
    restart: always
    logging:
      driver: "json-file"
      options:
        max-size: "10m"

Notice that we bind the port to 127.0.0.1:8080. This restricts external internet exposure, ensuring that only internal services (our Perplexica container) can query your search aggregator.

---

Step 3: Installing and Configuring Perplexica

With our data layer operational, navigate back to the root project folder to initialize the Perplexica orchestration layer:

cd ~/ai-search
git clone [https://github.com/ItzCrazyK0/Perplexica.git](https://github.com/ItzCrazyK0/Perplexica.git) frontend-engine
cd frontend-engine

Perplexica relies on a local configuration to link with your chosen LLM and backend search infrastructure. Duplicate the sample configuration file to begin editing:

cp sample.config.toml config.toml

Open config.toml and adjust the primary parameters to align with your API provider and the localized SearXNG service. If using Groq and our local SearXNG setup, your parameters should reflect the following:

[GENERAL]
PORT = 3000
SEARXNG_URL = "http://localhost:8080"

[API_KEYS]
GROQ = "your_groq_api_key_here"
OPENAI = ""

[MINION]
PROVIDER = "groq"
MODEL = "llama3-70b-8192"

To execute the deployment cleanly on our 2GB VPS architecture, utilize Perplexica's pre-built production Docker configurations. Run the deployment sequence in detached mode:

docker compose up -d
---

Performance Fine-Tuning and Production Considerations

Running an advanced stack on minimalist infrastructure requires continuous attention to optimization. Review these production-grade practices to maintain excellent response latency:

1. Cache Management

SearXNG can occasionally experience rate limits from upstream providers like Google or Bing if queried aggressively. Consider enabling a lightweight internal Redis cache layer within your SearXNG compose structure to temporarily store frequent queries and save outbound bandwidth.

2. Text Chunking & Embedding Constraints

Perplexica performs real-time text embedding to determine page relevance. Ensure you choose lightweight embedding models via the user interface settings (such as specialized HuggingFace API endpoints or OpenAI's text-embedding-3-small) to avoid spikes in local CPU utilization on your VPS.

3. Firewall Security

Ensure that ports 8080 (SearXNG backend) and 3000 (Perplexica internal backend) are closed to the public internet via your cloud provider's firewall dashboard. Only expose the frontend web interface, or protect the entire ecosystem behind a reverse proxy like Nginx or Caddy paired with basic HTTP Authentication for verified user access.

---

Conclusion: Sovereign AI Search is Within Reach

By combining Perplexica and SearXNG, you effectively decouple the power of modern conversational AI from intrusive, data-harvesting corporate monopolies. Offloading deep learning computation to specialized APIs while keeping data aggregation, synthesis logic, and tracking filtration completely localized allows you to run a cutting-edge AI engine flawlessly on a standard 2GB RAM VPS instance.

You now possess a fully sovereign, private, and customizable search assistant tailored to your business intelligence needs—built entirely on a highly optimized, cost-efficient infrastructure footprint.

Building a Self-Hosted AI Search Engine: Deploying Perplexica and SearXNG on a Budget 2GB VPS | DPTCloud