Back to articles
Technology Insight

Building an Automated 'AI Dev Swarm' to Resolve GitHub Issues: Integrating Cline, Tabby, and Ollama on a VPS

June 5, 2026

Introduction: The Era of Autonomous AI Development Swarms

The landscape of software engineering is undergoing a tectonic shift. We are moving rapidly from basic AI code-completion tools to fully autonomous multi-agent systems—often referred to as AI Dev Swarms. These systems don't just suggest the next line of code; they reason, navigate complex codebases, execute tests, and autonomously resolve real-world software defects.

For enterprise teams, deploying these autonomous agents often introduces a critical challenge: data privacy and infrastructure costs. Sending proprietary codebases to third-party commercial LLM providers can violate compliance policies and incur unpredictable API fees. The solution lies in building an independent, self-hosted AI development pipeline. This technical guide will walk you through setting up a local AI Dev Swarm capable of automatically fixing GitHub Issues by connecting Cline (the agentic orchestration engine), Tabby (the self-hosted code intelligence server), and Ollama (the local LLM runner) on a Virtual Private Server (VPS).

1. Understanding the Core Components

Before diving into the configuration, it is essential to understand the unique role each component plays within our automated development ecosystem:

  • Cline (formerly Devins Swarm/Roo Cline): An advanced, agentic framework that operates inside the development environment. Unlike passive chat assistants, Cline can read/write files, execute terminal commands, view compiler errors, and systematically iterate until a specific goal (like a GitHub issue) is resolved.
  • Tabby: An open-source, self-hosted AI coding assistant server. It acts as the centralized code intelligence repository, indexing your local codebase to provide fast, context-aware semantic search and code completion capabilities to the swarm.
  • Ollama: The heavy-lifter of the local AI stack. Ollama allows you to run powerful open-source Large Language Models (LLMs) such as Llama 3, Qwen2.5-Coder, or Mistral directly on your own hardware, providing the underlying reasoning engine for both Cline and Tabby.

2. Prerequisites and VPS Hardware Sizing

Running local LLMs and code indexers requires robust infrastructure. To host this swarm effectively on a VPS, we recommend the following minimum hardware specifications:

  • CPU: Minimum 8 vCPUs (optimized for compute-heavy workloads).
  • RAM: At least 32 GB RAM (essential for hosting 14B or 32B parameter models in memory).
  • Storage: 100 GB+ NVMe SSD (to store models and local code indexes).
  • GPU (Optional but highly recommended): An NVIDIA VPS instance (e.g., Tesla T4, A10G, or RTX 4090 hosted equivalents) will drastically accelerate inference speeds. If running CPU-only, expect slower execution times and use highly quantized models.
  • OS: Ubuntu 22.04 LTS or newer.

3. Step-by-Step Deployment Guide

Step 3.1: Setting up Ollama and the LLM Architecture

First, connect to your VPS via SSH and install Ollama. Run the following command in your terminal:

curl -fsSL [https://ollama.com/install.sh](https://ollama.com/install.sh) | sh

Once installed, we need to pull a model optimized specifically for coding tasks and agentic reasoning. Qwen2.5-Coder:14b or DeepSeek-Coder:base are exceptional open-source choices for this setup. Pull the model using:

ollama run qwen2.5-coder:14b

To ensure Cline and Tabby can communicate with Ollama remotely or across containerized environments, modify the Ollama service configurations to expose the host port by setting OLLAMA_HOST=0.0.0.0 in its environment variables, then restart the service.

Step 3.2: Deploying Tabby for Code Intelligence

Tabby will serve as our code repository indexer. The cleanest way to deploy Tabby on your VPS is using Docker. Create a docker-compose.yml file:

version: '3.8'
services:
  tabby:
    image: tabbyml/tabby
    ports:
      - "8080:8080"
    volumes:
      - ~/.tabby:/data
    command: serve --model Qwen2.5-Coder-1.5B --device cuda

Run docker compose up -d to start the server. Access the Tabby UI via http://your-vps-ip:8080, connect your target GitHub repositories, and let Tabby complete its structural indexing process. This ensures the AI swarm understands the structural relationships within your codebase.

Step 3.3: Configuring Cline as the Autonomous Swarm Agent

Cline operates natively as a headless CLI tool or via IDE extensions on headless environments. In your automation workflow on the VPS, configure Cline's provider profile to target your local setup. Create or edit the Cline configuration file (~/.config/cline/config.json):

{
  "apiProvider": "ollama",
  "ollama": {
    "baseUrl": "http://localhost:11434",
    "model": "qwen2.5-coder:14b"
  },
  "codeIntelligence": {
    "provider": "tabby",
    "endpoint": "http://localhost:8080"
  }
}

4. Automating the GitHub Issue Resolution Pipeline

To tie everything together into a hands-free automated pipeline, we utilize a lightweight automation script or a self-hosted CI/CD runner (like GitHub Actions Runner or GitLab Runner) deployed on the VPS. The workflow functions as follows:

  1. WebHook Trigger: A GitHub Webhook fires whenever a new issue with a specific label (e.g., bug:ai-fix) is created.
  2. Workspace Preparation: The automation script clones the repository into an isolated, sandboxed directory on the VPS.
  3. Swarm Activation: The script invokes Cline via the command line, passing the raw markdown text of the GitHub Issue as the primary objective.
Example Workflow Trigger Command:
cline --prompt "Fix the bug detailed in GitHub Issue #42: The payment gateway throws a null pointer exception when processing refunds. Use Tabby to search for the refund controller and apply the fix. Run 'npm test' to verify your solution before completing."

Upon invocation, Cline initiates an iterative loop. It asks Tabby for relevant code context, reads the faulty files, modifies the source code, runs localized tests on the VPS terminal, and inspects the errors if the tests fail. It repeats this loop autonomously until the test suite passes perfectly.

5. Closing the Loop: Automated Pull Requests

Once Cline satisfies the objectives and verifies the fix via testing, the automation wrapper executes the final phase of the pipeline. It switches to a new Git branch, commits the modified files, pushes the code back to the remote repository, and calls the GitHub API to generate a Pull Request.

The PR description is automatically populated with Cline's execution log, summarizing exactly which lines of code were modified and why. The human engineering team is left with a single task: reviewing and merging a pre-tested, fully formed fix.

6. Best Practices for Production Stability

Operating an autonomous AI swarm on private infrastructure requires rigorous safety guards. Adhere to these principles to maintain security and reliability:

  • Sandbox Execution: Never allow Cline to run commands directly on your bare metal VPS OS. Always isolate the execution context within ephemeral Docker containers to prevent accidental file deletion or malicious command execution.
  • Model Selection Maturation: For complex, multi-file refactoring, use larger quantized models (e.g., 32B or 70B parameters). Smaller models (7B or 14B) are highly efficient for straightforward syntax fixes but can break context chains during prolonged multi-turn reasoning.
  • Strict Prompt Boundaries: Implement system prompts that forbid the AI from altering environment files (.env), security keys, or CI/CD deployment configurations.

Conclusion

By connecting Cline, Tabby, and Ollama on a centralized VPS, you effectively build a private, zero-marginal-cost software engineering assistant that respects your data sovereignty. This setup shifts your human development team from manual firefighting to high-level system architecture, leveraging automated AI Dev Swarms to handle mundane issue remediation safely, scalably, and completely localized.