Back to articles
Technology Insight

Self-Hosting DeepSeek-R1 with Ollama and Open-WebUI on Low-Cost GPU Spot VPS

June 2, 2026

Introduction: The Open-Source AI Paradigm Shift

The landscape of artificial intelligence underwent a massive transformation with the release of DeepSeek-R1. As a first-tier reasoning model, DeepSeek-R1 matches or exceeds proprietary counterparts in complex math, coding, and logical reasoning tasks. However, relying on third-party APIs presents recurring costs, data privacy vulnerabilities, and rate-limiting bottlenecks for modern enterprises.

The alternative? Self-hosting. While local hardware demands can be prohibitive, leveraging GPU Spot instances (ephemeral, heavily discounted cloud GPUs) combined with lightweight orchestration tools like Ollama and Open-WebUI offers a production-grade solution. This architecture delivers total data sovereignty and elite reasoning performance at a fraction of standard cloud costs. This comprehensive guide walks you through the end-to-end deployment process.

Why Choose This Architecture? Efficiency Meets Cost-Optimization

Building a self-hosted AI stack requires balancing compute efficiency, cost, and user experience. Our curated stack addresses each vector perfectly:

  • DeepSeek-R1: Offers state-of-the-art reasoning capabilities. The distilled variants (such as the 8B, 14B, or 32B models based on Llama/Qwen) allow developers to choose the exact parameter size that fits their specific hardware budget.
  • Ollama: Serves as a highly optimized, lightweight execution engine that manages LLM weights, handles quantization smoothly, and exposes clean local inference endpoints.
  • Open-WebUI: Provides a polished, feature-rich graphical user interface mirroring ChatGPT, equipped with multi-user authentication, chat histories, RAG integration, and direct API routing.
  • GPU Spot Instances: Slashed infrastructure costs. Platforms like Vast.ai, RunPod, TensorDock, or spot offerings from AWS/GCP provide access to enterprise GPUs (like the RTX 3090, RTX 4090, or A100) at 70% to 90% discounts compared to on-demand pricing.

Step 1: Selecting and Provisioning Your Spot GPU VPS

Before executing code, you must select a hardware tier capable of loading DeepSeek-R1 weights into VRAM (Video RAM). The model size you choose dictates your GPU requirements:

  • DeepSeek-R1-Distill-Qwen-14B: Requires at least 12GB to 16GB VRAM (Recommended: 1x RTX 3090 or RTX 4090).
  • DeepSeek-R1-Distill-Qwen-32B: Requires at least 24GB to 32GB VRAM (Recommended: 1x RTX 4090 or 1x A6000).
  • DeepSeek-R1 (Full 671B): Requires massive distributed clusters (8x H100/A100). For cost-effective VPS setups, we strongly recommend focusing on the highly capable 14B or 32B distilled variants.

When deploying your instance on a platform like RunPod or Vast.ai, select a standard Ubuntu 22.04 LTS template with NVIDIA CUDA drivers pre-installed. Ensure you expose the following ports in your firewall settings:

  1. 11434 (Default port for Ollama API)
  2. 3000 (Default port for Open-WebUI)

Step 2: Installing Ollama and Downloading DeepSeek-R1

Once connected to your GPU Spot instance via SSH, update your system packages and install Ollama utilizing their official automated deployment script:

sudo apt update && sudo apt upgrade -y
curl -fsSL https://ollama.com/install.sh | sh

Verify that Ollama recognizes your GPU hardware by running the status command or checking the logs. Once verified, initiate the download and background execution of your selected DeepSeek-R1 model variant:

# To run the highly efficient 14B Qwen Distill variant
ollama run deepseek-r1:14b
Note: The download may take several minutes depending on your VPS network bandwidth. Once the download finishes, you can run a quick test prompt directly inside your terminal to confirm the model is outputting its signature tags correctly. Type /exit to return to the main shell.

Step 3: Deploying Open-WebUI via Docker

To provide an intuitive front-end interface for end-users, we deploy Open-WebUI. Using Docker simplifies dependency management and ensures isolated runtime execution. If Docker is not yet installed on your machine, install it using the following commands:

sudo apt install docker.io -y
sudo systemctl start docker
sudo systemctl enable docker

Next, spin up the Open-WebUI container. We will use the deployment configuration that links directly into the local host network where Ollama is actively listening:

sudo docker run -d -p 3000:8080 \
  -v open-webui:/app/backend/data \
  --name open-webui \
  --restart always \
  -e OLLAMA_BASE_URL=http://host.docker.internal:11434 \
  --add-host=host.docker.internal:host-gateway \
  ghcr.io/open-webui/open-webui:main

Give the container a few seconds to initialize its database. Open your web browser and navigate to http://your-vps-ip:3000. You will be greeted by the Open-WebUI account creation screen. The first account registered automatically becomes the Administrator account, allowing you to lock down registrations and manage model access control.

Step 4: Architecting for the Volatility of Spot Instances

While Spot instances offer unmatched cost efficiencies, they introduce a distinct risk: preemption. Cloud providers can reclaim Spot instances at any moment when demand surges. To protect your data and minimize downtime, you must build resilience into your deployment:

  • Externalize Data Volumes: Always mount your Open-WebUI database and Ollama models onto persistent network storage volumes separate from the root compute instance. If your container is terminated, you can attach the storage volume to a fresh GPU instance instantly.
  • Automated Backup Scripts: Configure a cron job that exports your Open-WebUI SQLite/PostgreSQL database to secure external object storage (such as AWS S3 or Cloudflare R2) every hour.
  • Infrastructure as Code (IaC): Maintain a clean initialization bash script containing the docker and tool installation steps outlined above. If your node gets terminated, automated initialization scripts can bring up an identical instance within two minutes.

Conclusion: Empowering Your Business with Sovereign AI

By coupling the logical reasoning performance of DeepSeek-R1 with the accessibility of Ollama and Open-WebUI, you establish a highly professional, secure, and responsive AI workspace. Shifting this infrastructure onto cheap GPU Spot VPS instances effectively eliminates the high premiums associated with proprietary SaaS offerings.

Whether you are analyzing confidential corporate datasets, generating secure source code, or prototyping custom agent workflows, this self-hosted architectural pattern ensures you retain complete command of your data pipeline, performance metrics, and infrastructure budget.

Self-Hosting DeepSeek-R1 with Ollama and Open-WebUI on Low-Cost GPU Spot VPS | DPTCloud