Back to articles
Technology Insight

Building a Self-Hosted AI Coding Server: Deploying Open WebUI and CodeGemma on a Linux VPS

May 27, 2026

Introduction: The Case for a Self-Hosted AI Coding Server

In the modern software development landscape, artificial intelligence has transitioned from a luxury to an absolute necessity. Tools like GitHub Copilot and ChatGPT have fundamentally changed how engineers write, debug, and document code. However, relying entirely on commercial, cloud-based AI providers introduces significant challenges for enterprises and independent developers alike. Key concerns include data privacy vulnerabilities, proprietary code leakage, potential intellectual property issues, and mounting monthly subscription costs.

The solution? Building your own private AI Coding Server. By leveraging a Linux Virtual Private Server (VPS), Google’s specialized CodeGemma large language model, and the intuitive Open WebUI interface, you can deploy a robust, self-hosted alternative. This setup ensures that your source code never leaves your infrastructure while delivering developer-focused AI assistance tailored to your workflow.

Prerequisites and System Architecture

Before diving into the deployment process, it is crucial to understand the architectural components and ensure your underlying infrastructure meets the heavy computational demands of running local large language models (LLMs).

Hardware Requirements

Running LLMs requires substantial memory and processing power. Depending on your project scale and budget, you can opt for either a CPU-optimized or GPU-accelerated Linux VPS:

  • CPU-Only Setup (Minimum): 4 vCPUs, 8GB RAM, and 50GB SSD. While slower, this is highly cost-effective for smaller models using quantization techniques.
  • GPU-Accelerated Setup (Recommended): Dedicated GPU with at least 8GB or 16GB VRAM (e.g., NVIDIA T4, A10G), 16GB System RAM, and 100GB NVMe SSD. This configuration guarantees near-instantaneous token generation and code completion.

Software Stack Component Overview

Our self-hosted environment relies on three primary building blocks working in harmony:

  1. Ollama: A lightweight, highly efficient framework designed for serving large language models locally on Linux environments.
  2. CodeGemma (by Google): A family of lightweight, powerful models built on top of Gemma, fine-tuned specifically for code-completion, code generation, and mathematical reasoning tasks.
  3. Open WebUI: An advanced, feature-rich user interface that mimics premium ChatGPT-like environments, featuring seamless integration with Ollama, multi-user management, and chat histories.

Step 1: Preparing Your Linux VPS Environment

To begin, connect to your Linux VPS via SSH and ensure all system packages are fully updated. For the purpose of this guide, we will use Ubuntu LTS as our base operating system.

sudo apt update && sudo apt upgrade -y

Next, install essential utilities and dependencies such as Docker, which will be required to run the Open WebUI interface efficiently and in isolation.

sudo apt install curl git docker.io docker-compose -y
sudo systemctl enable --now docker
Security Note: Always configure your system firewall (UFW) to restrict access to your AI server's ports. Ensure that only trusted IP addresses can communicate with your development backend.

Step 2: Installing Ollama and Fetching CodeGemma

Ollama simplifies the process of running models locally. Execute the official installation script directly onto your VPS:

curl -fsSL [https://ollama.com/install.sh](https://ollama.com/install.sh) | sh

Once the installation script completes, verified by checking the service status via systemctl status ollama, it is time to download Google's CodeGemma model. CodeGemma is available in multiple sizes. For standard VPS setups, the 7b (7 Billion parameter) instruct model offers an excellent balance between structural coding accuracy and speed.

ollama run codegemma:7b-instruct

This command automatically pulls the model weights from the official repository and launches an interactive terminal session. You can test its responsiveness by asking a simple question, such as: "Write a Python function to check for palindromes." Once verified, exit the interactive prompt using /exit. The Ollama service will continue running quietly in the background.

Step 3: Deploying Open WebUI with Docker

With the AI engine running seamlessly via Ollama, we need an accessible, elegant front-end. Open WebUI provides exactly that, complete with syntax highlighting, markdown rendering, and responsive design.

If Ollama and Open WebUI reside on the same Linux server, run the following Docker command to spin up the web interface, mapping port 8080 of the container to port 3000 of your host machine:

docker run -d -p 3000:8080 -v open-webui:/app/backend/data --add-host=host.docker.internal:host-gateway --name open-webui --restart always ghcr.io/open-webui/open-webui:main

The argument --add-host=host.docker.internal:host-gateway is critical; it allows the containerized Open WebUI instance to securely communicate with the host-bound Ollama service API.

Step 4: Configuring and Accessing Your AI Coding Server

Open your web browser and navigate to http://your-vps-ip:3000. On your first visit, you will be prompted to register an administrator account. The first account created holds root admin rights, allowing you to manage model settings, change UI themes, and control public registration behaviors to keep external users out.

After logging in, complete the configuration by following these steps:

  1. Navigate to Settings > Connections.
  2. Ensure the Ollama API URL point is correctly identified as [http://host.docker.internal:11434](http://host.docker.internal:11434).
  3. Save the configuration. Open WebUI will instantly pull available models.
  4. From the top model dropdown menu on the main dashboard, select codegemma:7b-instruct as your default model.

Maximizing CodeGemma for Daily Development

Your AI server is now fully operational. CodeGemma excels at specific engineering workflows that you can utilize directly through the Open WebUI chat interface:

  • Code Refactoring: Paste legacy code segments and instruct the model to optimize execution time, convert loops into list comprehensions, or reduce memory footprints.
  • Automated Unit Testing: Provide a function and request: "Generate comprehensive unit tests using pytest, covering edge cases and invalid inputs."
  • Architectural Advisory: Use CodeGemma to map out data models, design RESTful API structures, or debug obscure runtime errors from stack traces.

Conclusion: Freedom and Security in AI

By establishing a personal AI Coding Server using Open WebUI and CodeGemma running on a Linux VPS, you successfully break free from vendor lock-in and monthly recurring costs. More importantly, your intellectual property remains strictly confidential within an infrastructure you own and control. As open-source models continue to advance rapidly, your self-hosted architecture is primed to swap in newer, more capable models with a single terminal command, keeping you at the absolute cutting edge of AI-assisted engineering.

Building a Self-Hosted AI Coding Server: Deploying Open WebUI and CodeGemma on a Linux VPS | DPTCloud