Back to articles
Technology Insight

Building a Real-Time AI Voice Agent on a VPS: A High-Performance Open-Source Alternative to Vapi

June 1, 2026

Introduction: The Rise of Real-Time AI Voice Agents

In the rapidly evolving landscape of conversational artificial intelligence, latency is the ultimate user experience killer. When a customer speaks to an AI agent, they expect a fluid, human-like conversation. Traditional, fragmented systems that rely on separate steps for speech-to-text (STT), large language model (LLM) inference, and text-to-speech (TTS) synthesis often introduce frustrating delays of several seconds.

Platforms like Vapi and Bland AI have emerged to solve this problem, offering managed infrastructure for ultra-low latency voice bots. However, for enterprises requiring strict data privacy, full architectural control, and predictable scaling costs, relying on proprietary APIs can be a bottleneck. This technical guide explores how to architect, develop, and deploy a production-grade, real-time AI voice agent on a Virtual Private Server (VPS) using a powerful open-source stack: Pipecat and LiveKit.

The Architecture: Why Pipecat and LiveKit?

To build an alternative to Vapi, we need an ecosystem that handles WebRTC streaming, pipeline orchestration, and state management concurrently. Our open-source stack breaks down into two core layers:

  • LiveKit (The Transport Layer): A high-performance, WebRTC-based infrastructure designed for real-time audio and video streaming. LiveKit manages the bidirectional media pipeline between the user's browser or phone and our backend server, handling network jitter, packet loss, and disconnection gracefully.
  • Pipecat (The Orchestration Framework): Developed by Daily, Pipecat is an open-source framework specifically built for managing the complex state machine of AI voice agents. It acts as the glue that streams audio from LiveKit into an STT engine, pipes the transcribed text into an LLM, handles interruptions dynamically, and routes the generated text to an ultra-fast TTS provider before sending the audio back to LiveKit.

By bypassing closed ecosystems, this combination provides sub-second response times while running on your own independent infrastructure.

Prerequisites and Environment Setup

Before initiating development, ensure your deployment environment meets the following specifications:

1. VPS Provisioning

While an AI voice agent can run on CPU-only instances if you rely on external cloud APIs for LLM inference (e.g., OpenAI, Anthropic), your VPS must have sufficient throughput to handle multiple simultaneous WebRTC connections. We recommend a minimum configuration of:

  • 4 vCPUs
  • 8 GB RAM
  • Ubuntu 22.04 LTS or 24.04 LTS
  • A static public IP address with open ports for WebRTC (specifically HTTP/HTTPS ports and UDP ports for media streaming).

2. API Credentials

To power the intellectual and sensory capabilities of the agent, you will need access keys from the following providers:

  • Deepgram: For ultra-low latency, streaming Speech-to-Text.
  • OpenAI or Groq: For fast, conversational LLM responses (e.g., Llama-3 on Groq or GPT-4o-mini).
  • Cartesia or ElevenLabs: For highly realistic, streaming Text-to-Speech generation.

Step-by-Step Implementation

Let's walk through setting up the server component using Python, Pipecat, and the LiveKit SDK.

Step 1: Installing Dependencies

First, SSH into your VPS, set up a virtual environment, and install the necessary libraries. Pipecat offers modular plugins for different services, allowing you to swap out components easily.

pip install pipecat-ai[livekit,deepgram,openai,cartesia] python-dotenv asyncio

Step 2: Configuring LiveKit Server

You can run a LiveKit server instance directly on your VPS via Docker. Create a livekit.yaml configuration file and launch the container:

docker run --rm -d --name livekit-server \
  -p 7880:7880 -p 7881:7881 -p 50000-60000:50000-60000/udp \
  -v ./livekit.yaml:/livekit.yaml \
  livekit/livekit-server --config /livekit.yaml
Security Note: Ensure your firewall configuration allows incoming traffic on UDP ports 50000-60000 to enable peer-to-peer WebRTC media negotiations.

Step 3: Writing the Voice Agent Script

The core of our application is an asynchronous Python script that listens for incoming LiveKit room connections, instantiates the Pipecat pipeline, and starts processing media.

Create a file named agent.py and structure your Pipecat runner:

import asyncio
import os
from dotenv import load_dotenv
from pipecat.frames.frames import EndFrame
from pipecat.pipeline.pipeline import Pipeline
from pipecat.pipeline.runner import PipelineRunner
from pipecat.pipeline.task import PipelineTask
from pipecat.services.cartesia import CartesiaTTSService
from pipecat.services.deepgram import DeepgramSTTService
from pipecat.services.openai import OpenAILLMService
from pipecat.transports.services.livekit import LiveKitTransport

load_dotenv()

async def main():
    # Initialize transport layer with LiveKit
    transport = LiveKitTransport(
        url=os.getenv("LIVEKIT_URL"),
        token=os.getenv("LIVEKIT_TOKEN"),
    )

    # Initialize AI services
    stt = DeepgramSTTService(api_key=os.getenv("DEEPGRAM_API_KEY"))
    llm = OpenAILLMService(api_key=os.getenv("OPENAI_API_KEY"), model="gpt-4o-mini")
    tts = CartesiaTTSService(
        api_key=os.getenv("CARTESIA_API_KEY"),
        voice_id="sonic-english-male"
    )

    # Assemble the sequential processing pipeline
    pipeline = Pipeline([ 
        transport.input(),     # Capture user audio
        stt,                   # Audio -> Text
        llm,                   # Text -> Response Text
        tts,                   # Response Text -> Synthetic Audio
        transport.output()     # Play audio back to user
    ])

    task = PipelineTask(pipeline)
    runner = PipelineRunner()
    
    @transport.event_handler("on_participant_connected")
    async def on_participant_connected(transport, participant):
        # Greet the user when they enter the room
        await task.queue_frame(EndFrame()) # Clear buffers
        print(f"Participant {participant.identity} connected. Starting interaction.")

    await runner.run(task)

if __name__ == "__main__":
    asyncio.run(main())

Crucial Optimizations for VPS Production Deployments

Running real-time voice infrastructure requires optimizations that standard web applications rarely consider. To scale this setup effectively, keep these architectural practices in mind:

1. Handling User Interruption

One of the hardest elements of voice AI is managing when a user talks over the AI agent. Pipecat handles this natively via its frame-management architecture. When the STT engine detects speech from the user *while* the TTS service is outputting audio, Pipecat immediately fires a clearing frame down the pipeline, canceling the current TTS generation and flushing the audio playback queue. This gives the user the natural impression that the agent "stopped speaking to listen."

2. Minimizing Network Hop Latency

Geographical proximity is vital. If your users are primarily located in Southeast Asia, hosting your VPS in a European or US data center will append 200–300 milliseconds of network transit time to every exchange. Deploy your VPS in a regional hub close to your audience (e.g., Singapore or Tokyo) to ensure edge networks maintain minimal ping values.

3. Process Monitoring and Scaling

Because WebRTC connections are long-lived and stateful, traditional stateless load balancing (like simple round-robin HTTP configurations) does not work out of the box. Use process managers like PM2 or systemd units to ensure your python agent automatically recovers from unexpected runtime errors. For scaling, consider deploying a fleet of agent instances coordinated by LiveKit's Redis-backed multi-node architecture.

Conclusion and Next Steps

By leveraging Pipecat and LiveKit on an isolated VPS, you successfully build an enterprise-ready, ultra-low latency voice agent that functions as a highly cost-effective, transparent alternative to closed ecosystems like Vapi. You retain absolute ownership of the pipeline data, enjoy zero vendor markups on API usage, and can easily extend functionalities to support local LLMs or custom telephone network (SIP/PSTN) integrations.

As next steps, try implementing multi-turn memory buffers using Redis or integrating vector databases to give your real-time agent access to internal company documentation during live calls.

Building a Real-Time AI Voice Agent on a VPS: A High-Performance Open-Source Alternative to Vapi | DPTCloud