Back to articles
Technology Insight

Building a 24/7 Autonomous AI Coding Agent on VPS: Seamless Bug-Fixing with Gitness and Qwen2.5-Coder

June 2, 2026

Introduction to Autonomous Engineering

In the rapidly evolving landscape of software development, the integration of Artificial Intelligence has shifted from basic autocomplete tools to fully autonomous agents. While tools like GitHub Copilot enhance real-time writing, the true paradigm shift lies in autonomous background agents. Imagine a dedicated digital engineer operating 24/7 on your Virtual Private Server (VPS), monitoring repositories, analyzing failures, and proactively committing bug fixes while your core team sleeps.

This comprehensive guide explores how to architect and deploy a production-ready AI Coding Agent. By leveraging Gitness—an open-source, lightweight Git management and continuous integration platform—alongside Alibaba's state-of-the-art Qwen2.5-Coder Large Language Model (LLM), we will establish a self-healing development pipeline that minimizes downtime and drastically accelerates issue resolution.

The Architectural Blueprint: Gitness, VPS, and Qwen2.5-Coder

To build an agent capable of continuous execution without incurring prohibitive cloud costs, we require a highly efficient, self-hosted stack. The architecture consists of three core pillars:

  • The Hosting Layer (VPS): A cost-effective Virtual Private Server (e.g., DigitalOcean, Hetzner, or AWS EC2) running a lightweight Linux distribution serves as the stable, 24/7 environment.
  • The Repository and CI/CD Engine (Gitness): Created by Harness, Gitness provides a unified platform for source code management and automated pipelines. Its minimal footprint makes it ideal for VPS deployment compared to heavier alternatives.
  • The Intelligence Engine (Qwen2.5-Coder): This open-source LLM rivals proprietary models in coding proficiency. Optimized for code generation, reasoning, and bug fixing, it serves as the brain of our automated agent.

Why Choose Qwen2.5-Coder for Self-Hosted AI Agents?

Operating an AI agent running 24/7 requires a balance between inference cost, speed, and accuracy. Proprietary APIs can quickly become prohibitively expensive under continuous execution loops. Qwen2.5-Coder, particularly the 7B and 14B parameter variants, can be easily quantized and run locally on a mid-tier VPS using inference engines like Ollama or vLLM. It possesses a deep understanding of complex syntax, contextual logic, and regression testing, making it exceptionally qualified for autonomous debugging tasks.

Phase 1: Setting Up the VPS and Gitness Infrastructure

Before deploying our agent, we must prepare the server environment and initialize our source control platform. Follow these systematic steps to establish your foundation:

Step 1: System Optimization

Connect to your Linux VPS via SSH and update the core system packages to ensure stability and security:

sudo apt update && sudo apt upgrade -y
sudo apt install curl git docker.io docker-compose -y

Step 2: Deploying Gitness via Docker

Gitness is designed to run seamlessly inside a Docker container. Create a dedicated directory and launch the service using the following configuration:Create a docker-compose.yml file:

version: '3.8'
services:
  gitness:
    image: harness/gitness:latest
    ports:
      - "3000:3000"
    volumes:
      - /var/run/docker.sock:/var/run/docker.sock
      - ./gitness_data:/data
    restart: always

Execute docker-compose up -d to start Gitness. You can now access your repository manager by navigating to http://your-vps-ip:3000 in your browser and setting up your admin account.

Phase 2: Integrating Ollama and Qwen2.5-Coder

With Gitness operational, we need to host our intelligence layer on the same VPS (or a connected GPU-enabled instance). We will utilize Ollama for efficient local LLM management.

Step 1: Installing Ollama

Run the official installation script on your server:

curl -fsSL [https://ollama.com/install.sh](https://ollama.com/install.sh) | sh

Step 2: Pulling the Qwen2.5-Coder Model

For a standard VPS with modern multi-core CPUs and sufficient RAM (16GB+), the 7B parameter version offers an optimal balance of speed and precision. Run the following command to download the model:

ollama run qwen2.5-coder:7b

To verify the model is accessible via an API endpoint, test it with a simple curl request to port 11434. The agent will communicate through this local API to perform analysis.

Phase 3: Building the Autonomous Agent Logic

The core of our 24/7 agent is a background daemon (written in Python or Node.js) that monitors Gitness webhooks. When a pipeline fails, or a specific "bug" label is applied to an issue, the agent springs into action.

The Self-Healing Loop: Triggered by Error → Context Gathering → LLM Analysis → Code Patch Generation → Automated Testing → Pull Request Creation.

The Orchestration Script

Below is a conceptual framework of how the Python-based agent processes a repository error using the Gitness API and Qwen2.5-Coder:

import requests

def fetch_failed_pipeline_logs(repo_id, build_id):
    # Fetches standard error logs from the Gitness API
    headers = {"Authorization": "Bearer YOUR_GITNESS_TOKEN"}
    response = requests.get(f"http://localhost:3000/api/v1/repos/{repo_id}/builds/{build_id}/logs", headers=headers)
    return response.json()['logs']

def consult_qwen_coder(error_log, source_code):
    # Formulates a structured prompt for Qwen2.5-Coder
    prompt = f"""You are an expert autonomous software engineer.
    Review this error log:
    {error_log}
    
    And the corresponding source code:
    {source_code}
    
    Provide ONLY the corrected code blocks enclosed in valid diff format."""
    
    payload = {"model": "qwen2.5-coder:7b", "prompt": prompt, "stream": False}
    response = requests.post("http://localhost:11434/api/generate", json=payload)
    return response.json()['response']

The agent autonomously checks out the failing branch, applies the generated diff patch, and triggers a Gitness pre-commit pipeline to ensure the patch doesn't introduce syntax errors.

Phase 4: Automating the Bug-Fixing Pipeline

To ensure this ecosystem operates fully unattended 24/7, we embed the agent into Gitness Webhooks and Pipelines.

  1. Webhook Configuration: Navigate to your Gitness repository settings and create a webhook listening for Branch Updates and Pipeline Failures. Direct these payloads to your Python agent's listener port.
  2. State Verification: The agent maintains a local database (like SQLite) to track active tickets, preventing recursive loops where an imperfect AI fix causes another failure loop.
  3. Human-in-the-Loop Safeguards: While the agent operates autonomously, best practices dictate that it should not push directly to the main branch. Instead, configure the agent to push fixes to a structured branch (e.g., ai-fix/issue-104) and automatically open a Pull Request within Gitness, complete with an explanation of why the bug occurred and how it was resolved.

Best Practices for Production Deployment

Operating an autonomous agent requires strict boundaries to maintain security and resource efficiency. Adhere to these business guidelines:

  • Resource Throttling: Limit Docker and Ollama CPU usage via docker-compose resources constraints to prevent the AI model from consuming 100% of the VPS processing power during heavy inference cycles.
  • Scope Isolation: Give the Gitness API token granted to the AI agent restricted permissions. It should only have access to specific staging repositories, never core infrastructure configurations or master payment systems.
  • Prompt Engineering Refinement: Continually optimize the system prompt provided to Qwen2.5-Coder. Instruct it to adhere to your organization's specific formatting rules, linting configurations, and security practices (such as input sanitization).

Conclusion: The Future of DevOps

By marrying the lightweight efficiency of Gitness with the local processing power of Qwen2.5-Coder on a standard VPS, organizations can achieve a major competitive advantage. You effectively gain an tireless, unpaid junior engineer that continuously screens your codebase, analyzes pipeline crashes, and drafts pull requests overnight. Embrace this automated workflow today to drastically lower your Mean Time to Resolution (MTTR) and free your engineering talent to focus on building innovative features rather than chasing regression bugs.

Building a 24/7 Autonomous AI Coding Agent on VPS: Seamless Bug-Fixing with Gitness and Qwen2.5-Coder | DPTCloud