Back to articles
Technology Insight

Building a Self-Hosted 'Serverless AI Agent Workers' Infrastructure on Budget VPS with Bun and Ollama

May 25, 2026

Introduction: The Cost of the AI Revolution

The enterprise landscape is undergoing a massive shift toward autonomous AI agents—specialized software entities capable of planning, executing tasks, and reasoning. However, as organizations scale these workflows, they hit a critical financial and architectural bottleneck: the skyrocketing cost of proprietary LLM APIs. Relying solely on external endpoints for high-volume, repetitive agent tasks can quickly become unsustainable, especially for startups and mid-sized enterprises.

Furthermore, data privacy regulations and latency overhead add complications to public cloud dependencies. What if you could build a lightweight, self-hosted, and serverless-like AI agent infrastructure that runs efficiently on a budget Virtual Private Server (VPS)? By combining the blazing-fast execution of Bun with the local LLM orchestration of Ollama, you can achieve exactly that. This guide provides a comprehensive blueprint for architecting a high-performance, cost-effective AI worker pool without breaking the bank.

The Core Architectural Stack

To mimic a serverless architecture on a fixed-resource VPS, we need a software stack optimized for ultra-low memory footprint, rapid cold starts, and efficient concurrency management. Our architecture relies on three pillars:

  • Bun Runtime: A modern, all-in-one JavaScript runtime designed for speed. Bun replaces Node.js, offering significantly faster HTTP server startup times and native TypeScript support without transpilation overhead, making it perfect for lightweight worker instances.
  • Ollama: A streamlined framework for running large language models (such as Llama 3, Mistral, or Phi-3) locally. Ollama manages model weights, quantizations, and GPU/CPU acceleration behind a clean, OpenAI-compatible REST API.
  • The Serverless Agent Worker Pattern: Instead of keeping heavy agent processes running indefinitely in memory, we use Bun to spin up lightweight HTTP endpoints or queue consumers that process tasks on-demand, interacting with Ollama dynamically and releasing system resources when idle.

Step-by-Step Implementation Blueprint

1. Setting Up the Base Environment

First, secure a VPS from a budget-friendly provider. While a dedicated GPU is ideal for LLMs, modern quantized models run surprisingly well on standard CPU cores, provided you have at least 8GB to 16GB of RAM. Log into your clean Ubuntu server and install the necessary runtimes:

# Install Bun runtime
curl -fsSL [https://bun.sh/install](https://bun.sh/install) | bash

# Install Ollama
curl -fsSL [https://ollama.com/install.sh](https://ollama.com/install.sh) | bash

Verify the installations by checking the versions of both tools via your terminal. Once confirmed, pull a highly efficient, instruction-tuned model optimized for agentic tasks, such as Mistral or Llama 3 (8B quantized version):

ollama run llama3:8b

2. Crafting the Lightweight Worker with Bun

Next, we build our core worker script using Bun's native, ultra-fast Bun.serve() API. This worker acts as the serverless execution layer, receiving tasks via HTTP, processing them using an agentic chain, and communicating with the local Ollama instance.

Create a directory named ai-worker and initialize a new file called index.ts:

import { serve } from "bun";

const OLLAMA_ENDPOINT = "[http://127.0.0.1:11434/api/generate](http://127.0.0.1:11434/api/generate)";

serve({
  port: 3000,
  async fetch(req) {
    if (req.method !== "POST") {
      return new Response("Method Not Allowed", { status: 405 });
    }

    try {
      const { task, context } = await req.json();
      
      // Formulate the system prompt for the agent
      const systemPrompt = "You are an autonomous AI worker agent. Analyze the input, formulate a structured execution plan, and provide a concise response.";
      
      const response = await fetch(OLLAMA_ENDPOINT, {
        method: "POST",
        headers: { "Content-Type": "application/json" },
        body: JSON.stringify({
          model: "llama3:8b",
          prompt: `Context: ${context}\nTask: ${task}`,
          system: systemPrompt,
          stream: false
        })
      });

      const data = await response.json();
      return new Response(JSON.stringify({ success: true, result: data.response }), {
        headers: { "Content-Type": "application/json" }
      });

    } catch (error) {
      return new Response(JSON.stringify({ success: false, error: error.message }), { status: 500 });
    }
  }
});
console.log("Serverless AI Worker listening on port 3000...");

3. Emulating Serverless Scaling on a Budget VPS

To simulate true serverless execution on a single VPS, we must ensure that our workers can scale concurrently under load but won't crash the system when resources get tight. We can achieve this by implementing a reverse proxy and process manager wrapper.

  • PM2 or Bun Cluster Mode: Utilize Bun's native clustering capabilities or use a process manager like PM2 to spawn multiple worker instances across your available CPU cores. This ensures that network I/O bound tasks do not block the primary thread.
  • Nginx Reverse Proxy: Place Nginx in front of your Bun workers. Nginx can handle incoming SSL termination, implement rate-limiting to prevent hardware starvation, and load-balance traffic across your internal Bun ports.
Architecture Note: Because Ollama processes LLM inferences sequentially or shifts them into internal queues based on configuration, your Bun workers act as excellent traffic buffers, protecting the underlying AI engine from memory over-allocation.

Performance Optimization Strategies for Budget Hardware

Running local AI agents on a budget VPS requires strict resource discipline. Implementing the following techniques will maximize your hardware's throughput:

Model Quantization

Never deploy unquantized full-precision models on standard VPS hardware. Always opt for 4-bit or 5-bit quantized models (often labeled as q4_K_M or q5_K_M within the Ollama repository). These models reduce memory consumption by up to 70% while retaining over 95% of the base model's reasoning capabilities.

Context Window Management

Large context windows consume exponential amounts of RAM. For specific agentic workers tasked with discrete operations (like data extraction or sentiment analysis), restrict the context window explicitly in your Ollama parameters to 2048 or 4096 tokens instead of letting it default to larger thresholds.

Response Streaming

If your AI workers communicate with a web UI or a continuous workflow manager, modify the Bun code to stream responses chunk-by-chunk rather than waiting for the complete JSON generation. This drastically lowers the perceived latency (Time to First Token) for end users.

Conclusion and Future Outlook

By marrying the performance-focused design of Bun with the local inference power of Ollama, you effectively democratize AI agent deployment. You no longer need thousands of dollars in cloud infrastructure to run private, secure, and highly responsive autonomous workers. A standard $20 to $40 per month VPS can successfully host a dedicated fleet of agents capable of handling background tasks, content generation, and data analysis around the clock.

As the open-source community continues to release smaller, more powerful models, the efficiency of this self-hosted serverless paradigm will only improve, giving forward-thinking companies a distinct competitive advantage in the AI era.