Back to articles
Technology Insight

Building a Sub-300ms Low-Latency AI Voice Agent Using Vocode, LiveKit, and Docker VPS

June 2, 2026

Introduction: The Race for Real-Time Conversational AI

In the landscape of modern customer engagement, latency is the ultimate dealbreaker. While text-based Generative AI has transformed business workflows, bringing the same fluid interaction to voice communication presents a massive technical hurdle. Traditional voice bots often suffer from awkward pauses, halting speech patterns, and delays that exceed two to three seconds. For a business, this translates to frustrated users and abandoned calls.

Achieving a human-like conversation requires a response time of under 300 milliseconds. Achieving this milestone requires an optimized, low-latency stack designed specifically for real-time audio streaming. In this technical guide, we will explore how to architect, deploy, and optimize an enterprise-grade AI Voice Agent utilizing Vocode for conversational orchestration, LiveKit for WebRTC audio transport, and Docker for scalable VPS deployment.

The Core Architecture: Why Vocode and LiveKit?

Building a real-time voice assistant requires orchestrating three fundamental pillars of speech processing simultaneously:

  • Automated Speech Recognition (ASR): Converting user audio into text instantly.
  • Large Language Model (LLM): Generating contextual text responses with streaming outputs.
  • Text-to-Speech (TTS): Synthesizing text back into natural human voice.

Managing these components independently over standard HTTP requests introduces severe latency. That is where Vocode and LiveKit come into play.

Vocode: The Orchestration Layer

Vocode acts as the central brain of our voice agent. It abstracts the complex state management required for bidirectional streaming conversations. Instead of waiting for a user to finish a whole sentence, Vocode processes chunks of audio sequentially, piping them directly into streaming ASR and LLM endpoints. It also handles crucial conversational nuances such as interuption handling, ensuring that if a user speaks while the bot is talking, the bot stops instantly to listen.

LiveKit: The WebRTC Powerhouse

Traditional WebSockets are often insufficient for high-fidelity, ultra-low-latency audio due to TCP head-of-line blocking. LiveKit solves this by utilizing WebRTC (UDP). It provides an enterprise-grade, open-source infrastructure designed for real-time video and audio transport. By implementing LiveKit, audio packets are delivered with minimal overhead, maintaining stability even under fluctuating network conditions.

Prerequisites and Environment Setup

Before initiating the deployment, ensure your virtual private server (VPS) meets the following baseline specifications:

  • OS: Ubuntu 22.04 LTS or newer
  • Hardware: Minimum 2 vCPUs, 4GB RAM (8GB recommended for production)
  • Network: Public IPv4 address with ports 443 (HTTPS), 7801 (LiveKit WebRTC), and 5000/TCP open.
  • Software: Docker Engine v20.10+ and Docker Compose v2.0+ installed.

Additionally, you will need API keys from ultra-low-latency providers: Deepgram (for Nova-2 ASR), OpenAI (for GPT-4o streaming), and ElevenLabs or Cartesia (for fast TTS generation).

Step-by-Step Deployment Guide on Docker VPS

Step 1: Configuring the LiveKit Server

First, we need to generate a secure configuration file for LiveKit. Create a directory named livekit and initialize the livekit.yaml configuration:

port: 7801
rtc:
udp_port_dev: 7801
use_external_ip: true
keys:
API_KEY_PLACEHOLDER: API_SECRET_PLACEHOLDER

Ensure you replace the placeholders with secure, auto-generated alphanumeric keys. This ensures that only authorized clients and your Vocode agent can access the WebRTC audio rooms.

Step 2: Crafting the Vocode Core Application

Our core application will bind Vocode’s environment to the LiveKit room. Below is the conceptual layout of our Dockerized Python application structure:

├── docker-compose.yml
├── livekit/
│   └── livekit.yaml
└── vocode-agent/
    ├── Dockerfile
    ├── main.py
    └── requirements.txt

In main.py, we initialize the LiveKit Room workflow and connect Vocode’s StreamingConversation class, passing the real-time configurations for Deepgram, OpenAI, and Cartesia. The agent is configured to listen continuously and stream responses byte-by-byte back to the LiveKit audio track.

Step 3: Orchestrating with Docker Compose

To tie both services together seamlessly, we utilize a unified docker-compose.yml file. This architecture ensures isolated networking and automated container restarts.

Note: Network performance is critical. We expose the host network or map specific UDP ports directly to bypass Docker network bridging overhead, shaving off vital milliseconds.

version: '3.8'
services:
  livekit:
    image: livekit/livekit-server:v1.5
    command: --config /livekit.yaml
    volumes:
      - ./livekit/livekit.yaml:/livekit.yaml
    network_mode: "host"
    restart: always

  vocode-agent:
    build: ./vocode-agent
    environment:
      - LIVEKIT_URL=http://localhost:7801
      - LIVEKIT_API_KEY=YOUR_LIVEKIT_KEY
      - LIVEKIT_API_SECRET=YOUR_LIVEKIT_SECRET
      - DEEPGRAM_API_KEY=YOUR_DEEPGRAM_KEY
      - OPENAI_API_KEY=YOUR_OPENAI_KEY
      - CARTESIA_API_KEY=YOUR_CARTESIA_KEY
    network_mode: "host"
    depends_on:
      - livekit
    restart: always

Optimizing for Sub-300ms Latency

Deploying the containers is only half the battle. To cross the sub-300ms threshold consistently, specific low-latency configurations must be strictly applied across your pipeline:

  1. Geographic Proximity: Deploy your VPS in a data center closest to your target audience. If your users are in Southeast Asia, hosting your VPS in Singapore or Tokyo reduces physical fiber transit delays by up to 100ms.
  2. Audio Chunk Size Optimization: Configure your ASR client to send audio in chunks of 20ms to 40ms. Larger chunks increase buffer latency, while smaller chunks risk packet fragmentation.
  3. LLM Time-to-First-Token (TTFT): Use highly optimized models like GPT-4o-mini or a fine-tuned Llama 3 model run locally via vLLM. Ensure streaming mode is enabled so that the TTS engine receives text tokens immediately as they are generated.
  4. WebSocket & WebRTC Tuning: Disable Nagle’s algorithm (TCP_NODELAY) on your application backend to force packets to send instantly without buffering.

Conclusion: The Future of Voice-First Business

By architectural integration of Vocode's orchestration and LiveKit’s real-time WebRTC infrastructure inside a optimized Docker VPS, businesses can break free from high-latency barriers. A sub-300ms AI Voice Agent feels natural, handles interruptions perfectly, and mimics human interaction quality flawlessly. Whether scaled for automated customer support, real-time translation, or voice commerce, this stack provides the robust, cost-effective foundation required for next-generation voice applications.

Building a Sub-300ms Low-Latency AI Voice Agent Using Vocode, LiveKit, and Docker VPS | DPTCloud