Back to articles
Technology Insight

Build Your Own Multi-Tasking AI Coding Agent on a 4GB RAM VPS: A Self-Hosted Alternative to Cursor and Bolt.new

May 26, 2026

Introduction

The developer landscape is undergoing a massive shift. Tools like Cursor, Devin, and Bolt.new have demonstrated the immense power of AI-driven development, allowing engineers to scaffold, debug, and deploy applications using natural language. However, relying on closed-source, SaaS-based AI coding assistants introduces significant challenges: soaring subscription costs, restrictive API rate limits, and severe data privacy concerns regarding proprietary source code.

What if you could build and control your own multi-tasking AI Coding Agent, hosted entirely on a budget-friendly Virtual Private Server (VPS) with only 4GB of RAM? Thanks to the optimization of open-source projects, this is no longer a theoretical concept. In this guide, we will walk through engineering a self-hosted, production-ready AI Coding Agent using Qwen-Agent as the orchestration framework and LiteLLM as the unified model proxy gateway. This setup offers a resilient, private, and highly cost-effective alternative to commercial platforms.

---

The Architectural Blueprint: Why Qwen-Agent and LiteLLM?

Operating an advanced AI agent on a constrained environment like a 4GB RAM VPS requires a highly efficient architectural stack. Trying to run a massive 70B parameter model locally on such hardware is impossible. Instead, our strategy leverages a hybrid edge-cloud architecture, utilizing highly efficient local orchestration coupled with optimized API endpoints.

1. Qwen-Agent: The Orchestration Engine

Developed by the Alibaba Qwen team, Qwen-Agent is a robust framework designed for building applications based on Large Language Models (LLMs). It natively supports Function Calling (Tool Use), Memory Management, and Multi-Agent Orchestration. Unlike heavier frameworks that consume excessive memory, Qwen-Agent is highly optimized, making it the perfect brain for our resource-constrained VPS.

2. LiteLLM: The Universal Proxy Gateway

Managing multiple AI providers, fallback strategies, and cost tracking can quickly become a DevOps nightmare. LiteLLM solves this by acting as a universal proxy. It translates standard OpenAI-formatted API calls into requests compatible with over 100 LLM providers (Anthropic, OpenAI, Groq, OpenRouter, or self-hosted Ollama instances). By running LiteLLM locally on our VPS, we gain a centralized abstraction layer to manage models seamlessly without changing a single line of application code.

---

Prerequisites and Environment Setup

Before proceeding with the installation, ensure your environment meets the following baseline requirements:

  • Hardware: A Linux VPS (Ubuntu 22.04 LTS or 24.04 LTS preferred) with at least 2 vCPUs and 4GB of RAM.
  • Prerequisites: Docker and Docker Compose installed, a non-root user with sudo privileges, and a registered domain name (optional, but highly recommended for secure HTTPS access).
  • API Keys: Access keys from your preferred LLM providers (e.g., Anthropic Claude 3.5 Sonnet for advanced coding, or Groq/OpenRouter for high-speed, low-cost operations).
---

Step-by-Step Implementation Guide

Step 1: System Optimization for 4GB RAM

When working with a 4GB RAM constraint, operating system optimization is critical to prevent the Linux kernel from killing processes due to Out-Of-Memory (OOM) errors. We will begin by configuring a robust Swap space.

# Create a 4GB swap file
sudo fallocate -l 4G /swapfile
sudo chmod 600 /swapfile
sudo mkswap /swapfile
sudo swapon /swapfile

# Make the swap permanent
echo '/swapfile none swap sw 0 0' | sudo tee -a /etc/fstab

# Adjust swappiness to utilize swap only when necessary
sudo sysctl vm.swappiness=10
echo 'vm.swappiness=10' | sudo tee -a /etc/sysctl.conf

Step 2: Deploying LiteLLM via Docker Compose

Next, we will deploy LiteLLM to handle our model routing. Create a dedicated project directory and define the configuration files.

Create a litellm-config.yaml file to define your model routing rules:

model_list:
  - model_name: coding-model
    litellm_params:
      model: claude-3-5-sonnet-20241022
      api_key: "os.environ/ANTHROPIC_API_KEY"
  - model_name: fast-fallback
    litellm_params:
      model: openai/gpt-4o-mini
      api_key: "os.environ/OPENAI_API_KEY"

Now, create the docker-compose.yml file to spin up the LiteLLM container service:

version: '3.8'
services:
  litellm:
    image: ghcr.io/berriai/litellm:main-latest
    ports:
      - "4000:4000"
    environment:
      - ANTHROPIC_API_KEY=your_actual_anthropic_key_here
      - OPENAI_API_KEY=your_actual_openai_key_here
      - LITELLM_MASTER_KEY=sk_live_vps_agent_secure_key_2026
    volumes:
      - ./litellm-config.yaml:/app/config.yaml
    command: ["--config", "/app/config.yaml", "--detailed_debug", "False"]
    restart: always

Run docker compose up -d to launch the proxy service. LiteLLM is now exposing an OpenAI-compatible endpoint at http://localhost:4000.

Step 3: Developing the Multi-Tasking Agent with Qwen-Agent

With our model gateway operational, we can now build the autonomous coding agent. Qwen-Agent allows us to equip our assistant with specialized tools such as file system manipulation, bash execution, and code analysis capabilities.

First, install the required python dependencies within a virtual environment:

python3 -m venv venv
source venv/bin/activate
pip install qwen-agent openai requests

Create our primary agent application script, named agent.py:

import os
from qwen_agent.agents import Assistant
from qwen_agent.tools.base import BaseTool, register_tool

# Define a custom tool for secure file execution/manipulation
@register_tool('file_writer')
class FileWriter(BaseTool):
    description = 'Writes or updates source code files in the workspace.'
    parameters = {
        'type': 'object',
        'properties': {
            'path': {'type': 'string', 'description': 'Relative path to the file'},
            'content': {'type': 'string', 'description': 'The exact code content to write'}
        },
        'required': ['path', 'content']
    }

    def call(self, params: str, **kwargs) -> str:
        params = self._verify_json_format(params)
        path = params['path']
        content = params['content']
        
        # Secure sandbox path restriction
        secure_path = os.path.abspath(os.path.join('./workspace', path))
        if not secure_path.startswith(os.path.abspath('./workspace')):
            return "Error: Security violation. Access outside workspace denied."
            
        os.makedirs(os.path.dirname(secure_path), exist_ok=True)
        with open(secure_path, 'w', encoding='utf-8') as f:
            f.write(content)
        return f"Successfully wrote code to {path}"

# Configure LLM client via our local LiteLLM Proxy Gateway
llm_cfg = {
    'model': 'coding-model',
    'model_server': 'http://localhost:4000/v1',
    'api_key': 'sk_live_vps_agent_secure_key_2026',
}

system_instruction = (
    "You are an elite autonomous multi-tasking AI Coding Agent. "
    "Your task is to plan, architect, write, and verify clean code based on user requirements. "
    "Always ensure proper error handling and prioritize modular structure."
)

bot = Assistant(llm=llm_cfg,
                system_instruction=system_instruction,
                function_list=['file_writer', 'code_interpreter'])

if __name__ == '__main__':
    print("AI Coding Agent Initialized. Enter your development request:")
    user_prompt = input("> ")
    
    messages = [{'role': 'user', 'content': user_prompt}]
    print("\n[Agent is planning and executing tools...]")
    
    for response in bot.run(messages=messages):
        if 'content' in response[-1]:
            print(response[-1]['content'])
---

Performance and Resource Optimization

Running this infrastructure on a 4GB RAM VPS requires keeping a close eye on resource consumption. To maximize efficiency, consider the following production best practices:

Pro Tip: Implement regular Docker log rotation. Left unchecked, container logs can grow rapidly and completely consume your disk space, leading to system failure.
  • Memory Cap Enforcement: Set explicit memory constraints in your Docker Compose files (e.g., mem_limit: 1g for LiteLLM) to prevent any single service from causing a system-wide crash.
  • Connection Pooling: Utilize LiteLLM's internal pooling mechanisms to minimize connection overhead and lower CPU usage during concurrent operations.
  • Model Fallbacks: Configure your routing rules to automatically transition to a lighter, faster model (like GPT-4o-mini or Mistral-7B via OpenRouter) if a primary model experiences high latency or timeout issues.
---

Conclusion

Building a self-hosted AI Coding Agent on a 4GB RAM VPS proves that you do not need expensive commercial subscriptions or heavy local hardware to leverage advanced AI workflows. By combining Qwen-Agent's efficient orchestration with LiteLLM's flexible routing, you gain complete ownership over your development workspace, robust code privacy, and full control over your operational costs.

As next steps, you can expand this system by exposing a web-based frontend using Streamlit or deploying an open-source terminal interface like Open WebUI, fully cementing your customized, self-hosted alternative to Cursor and Bolt.new.

Build Your Own Multi-Tasking AI Coding Agent on a 4GB RAM VPS: A Self-Hosted Alternative to Cursor and Bolt.new | DPTCloud