Building a 24/7 Autonomous AI Coding Agent on VPS with Gitness and LLM-Driven Self-Healing Workflows
Introduction: The Dawn of the Autonomous Developer
The paradigm of software development is undergoing a seismic shift. We are moving rapidly from AI-assisted coding—where developers use tools like Copilot to autocomplete lines of code—to autonomous AI agents capable of managing entire development lifecycles. Imagine an engineer that never sleeps, constantly monitoring your repositories, identifying bottlenecks, and autonomously writing, testing, and deploying bug fixes 24/7.
In this technical guide, we will explore how to build a production-grade, background-running AI Coding Agent deployed on a Virtual Private Server (VPS). By combining Gitness (an open-source, lightweight development platform that integrates source control management with pipelines) and advanced Large Language Models (LLMs), we will construct a self-healing workflow that automatically fixes bugs when tests fail. Here is how you can achieve complete automation for your development ecosystem.
---Why Gitness and a Self-Hosted VPS?
While commercial platforms offer robust CI/CD features, hosting your AI agent infrastructure on a dedicated VPS combined with Gitness offers three distinct enterprise advantages:
- Cost Efficiency: Running continuous LLM agents and pipeline loops on proprietary cloud platforms can incur unpredictable, exorbitant costs. A VPS provides predictable monthly budgeting.
- Data Privacy and Security: Codebases often contain proprietary logic and intellectual property. Keeping your code management (Gitness) and AI processing loop within a private VPS infrastructure ensures complete data sovereignty.
- Deep Pipeline Control: Gitness is incredibly lightweight, resource-efficient, and designed to run seamlessly in Docker environments, making it the perfect orchestrator for resource-constrained VPS instances.
Architectural Blueprint of the 24/7 AI Coding Agent
Before diving into the code and configuration, it is essential to understand how data and commands flow through our autonomous ecosystem. The architecture operates in a continuous, event-driven feedback loop:
- Code Trigger: A developer pushes code or an automated system triggers a cron-job test run in Gitness.
- Pipeline Failure: The Gitness pipeline executes the test suite. If a bug is present, the test fails, generating a detailed error log.
- Agent Activation: A failure webhook triggers our background AI Agent daemon, running continuously via a process manager like
PM2or Docker Compose. - Context Assembly: The agent fetches the failing code file, the associated test script, and the exact error stack trace from Gitness.
- LLM Inference & Self-Healing: The agent queries a specialized LLM with a structured prompt, requesting a precise code patch.
- Verification & PR Creation: The agent applies the patch locally, runs the tests again, and if successful, pushes a new branch and automatically opens a Pull Request (PR) in Gitness.
Note: By isolating the AI's changes to a separate Git branch and requiring a successful local test run, we ensure the agent can never break the main production branch.---
Step 1: Setting Up Gitness on Your VPS
First, we need to deploy Gitness on our VPS. Ensure your server has Docker and Docker Compose installed. Create a docker-compose.yml file to spin up the Gitness service:
version: '3.8'
services:
gitness:
image: harness/gitness:latest
ports:
- "3000:3000"
volumes:
- ./gitness_data:/data
restart: always
environment:
- GITNESS_DEBUG=trueRun docker compose up -d to initialize the platform. Navigate to http://your-vps-ip:3000, create your admin account, and initialize your project repository. Gitness provides an intuitive UI coupled with powerful underlying API endpoints that our agent will exploit.
Step 2: Designing the Gitness Pipeline for Error Capture
To enable self-healing, our pipeline must explicitly export logs when a failure occurs. In your repository root, create a .gitness.yaml file. This pipeline compiles the application and runs unit tests. Crucially, it includes a post-failure step that alerts our background agent.
pipeline:
- name: test-application
image: node:18
commands:
- npm install
- npm test
- name: trigger-ai-agent
image: curlimages/curl:latest
status: [ failure ]
commands:
- curl -X POST http://localhost:8080/webhook-fail -H "Content-Type: application/json" -d '{"repo": "'"$GITNESS_REPO_NAME"'", "commit": "'"$GITNESS_COMMIT_SHA"'"}'Using the status: [ failure ] directive ensures that the trigger-ai-agent step only executes if the application code contains bugs that break the test suite.
Step 3: Developing the Background AI Agent Daemon
The core of our automation is a Node.js or Python daemon running continuously on the VPS. This server listens for the failure webhook, leverages an LLM (such as GPT-4o or a self-hosted CodeLlama instance via Ollama) to diagnose the issue, and applies the fix.
Below is a conceptual implementation of the agent's core workflow written in Python:
from flask import Flask, request
import requests
import os
import openai
app = Flask(__name__)
@app.route('/webhook-fail', methods=['POST'])
def handle_failure():
data = request.json
repo_name = data['repo']
commit_sha = data['commit']
# 1. Fetch failing context from Gitness API
error_logs = fetch_gitness_pipeline_logs(repo_name, commit_sha)
source_code = fetch_repository_file(repo_name, "src/index.js")
# 2. Invoke LLM for Self-Healing
corrected_code = generate_llm_patch(source_code, error_logs)
# 3. Apply Patch and Validate
apply_patch_to_branch(repo_name, corrected_code)
return {"status": "Agent processed failure, patch pushed."}, 200
def generate_llm_patch(original_code, error_logs):
prompt = f"""
You are an expert autonomous software engineer.
The following source code is breaking during testing.
--- ORIGINAL CODE ---
{original_code}
--- ERROR LOGS ---
{error_logs}
Identify the bug and rewrite the complete file with the bug fixed.
Respond ONLY with the raw code inside standard markdown blocks.
"""
response = openai.ChatCompletion.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}]
)
return response.choices[0].message.contentTo guarantee this agent runs 24/7 without interruption, manage the process using PM2 (pm2 start agent.py --name "ai-agent") or encapsulate it within its own Docker container with a restart: always policy.
Step 4: Crafting the System Prompt for Precision Self-Healing
The primary point of failure for AI agents is hallucination or incomplete code generations. To make your agent highly reliable, you must enforce a strict response format using advanced prompting techniques like Role Prompting and Structured Output Constraints.
Your agent should explicitly instruct the LLM to:
- Analyze the stack trace to isolate the exact line causing the runtime error.
- Assess edge cases that could have triggered the specific test assertion failure.
- Return the entire file contents without shortening code with comments like "// rest of your code remains the same".
Monitoring, Observability, and Guardrails
Allowing an AI agent to continuously write code and trigger pipelines requires strict guardrails to prevent infinite loops (e.g., the AI introduces a new bug trying to fix an old one, causing endless failing pipelines). Implement the following safety checks:
- Loop Counter: Limit the agent to a maximum of 3 sequential fix attempts per branch. If the tests still fail after 3 attempts, the agent must halt and assign a human developer to the ticket.
- Strict Sandboxing: Ensure the background agent executes its verification tests inside an isolated Docker container, completely segregated from the host VPS file system to prevent accidental file deletion or security breaches.
Conclusion: The Future of DevOps
Deploying a 24/7 autonomous AI Coding Agent running seamlessly alongside Gitness changes the dynamics of continuous integration. By automating the tedious loop of identifying test failures, reading logs, and writing patches, human developers are freed to focus on architecture, product design, and strategic business logic. Setting up this framework on a private VPS ensures you maintain ultimate control over your intellectual property and operational costs while staying at the absolute cutting edge of software engineering automation.
