Back to articles
Technology Insight

Building a Real-Time AI Voice Agent with Under 300ms Latency Using Pipecat, Daily Transport, and Docker

June 2, 2026

Introduction: The Race to Ultra-Low Latency Voice AI

In the landscape of conversational AI, voice is the ultimate frontier. While Large Language Models (LLMs) have mastered text-based communication, transitioning that capability into real-time voice interactions introduces a formidable adversary: latency. In human conversation, the typical response gap is roughly 200 to 300 milliseconds. If an AI voice agent takes longer than 500ms to respond, the illusion breaks, resulting in awkward pauses, overlapping speech, and a frustrating user experience.

Building a system that ingests audio, transcribes it, processes it through an LLM, synthesizes a voice response, and streams it back to the user—all in under 300ms—requires a radical departure from traditional HTTP-based architectures. This technical guide explores how to build a production-grade, real-time AI Voice Agent using Pipecat, Daily Transport, and Docker to achieve sub-300ms latency.

The Core Architectural Stack

To achieve near-instantaneous responses, every layer of the conversational pipeline must be optimized for streaming. We move away from request-response paradigms toward a continuous, bidirectional pipeline. Our stack leverages three foundational technologies:

  • Pipecat: An open-source framework designed specifically for building voice and multimodal AI agents. Pipecat handles the complex orchestration of frame-based data, managing pipelines that connect audio input, Speech-to-Text (STT), LLMs, and Text-to-Speech (TTS) engines.
  • Daily Transport: Built on WebRTC (Web Real-Time Communication), Daily provides the ultra-low latency infrastructure required to transmit raw audio and video frames globally over UDP, bypassing the overhead of traditional TCP/HTTP streaming.
  • Docker: Containerization ensures consistency across development and production environments, allowing us to isolate dependencies, manage environment variables, and scale instances horizontally.

Understanding the 300ms Latency Budget

To achieve a sub-300ms total response time, we must enforce a strict latency budget across each component of the AI pipeline. A typical breakdown looks like this:

  1. Network Ingestion (WebRTC): ~20ms – Daily Transport delivers audio frames from the client to our backend via optimized WebRTC routing.
  2. Speech-to-Text (STT): ~50ms – Utilizing deep-gram streaming or local Whisper models with voice activity detection (VAD) to transcribe words the moment they are spoken.
  3. LLM Time-to-First-Token (TTFT): ~120ms – Leveraging fast, streaming-optimized models (like Groq, Together AI, or GPT-4o-mini) that emit tokens instantaneously.
  4. Text-to-Speech (TTS) Generation: ~80ms – Streaming chunk-based audio synthesis (using Cartesia or ElevenLabs Turbo) so audio playback begins before the LLM finishes its entire thought.
  5. Network Delivery: ~20ms – WebRTC streams the generated audio back to the client headset.
Total Estimated Budget: ~290ms. By managing data as continuous frames rather than whole sentences, the user hears the beginning of a response while the rest of the sentence is still being computed.

Step-by-Step Implementation

1. Setting Up the Pipecat Pipeline

The core of our agent relies on a Pipecat pipeline initialization. Unlike standard applications, a Pipecat agent operates as a state machine processing individual audio and text frames. Below is the conceptual architecture of our Python-based agent application:

import asyncio
from pipecat.frames.frames import EndFrame
from pipecat.pipeline.pipeline import Pipeline
from pipecat.pipeline.task import PipelineTask
from pipecat.services.openai import OpenAILLMService
from pipecat.services.cartesia import CartesiaTTSService
from pipecat.services.deepgram import DeepgramSTTService
from pipecat.transports.services.daily import DailyTransport

We initialize the DailyTransport, configuring it to handle bi-directional audio. Next, we instantiate our streaming services. Deepgram is selected for its ultra-fast streaming STT API, OpenAI (or Groq) handles the LLM intelligence, and Cartesia provides high-fidelity, sub-100ms voice synthesis.

We then orchestrate them into a unified pipeline where the output of the STT flows directly into the input of the LLM, and the token output of the LLM feeds directly into the TTS engine, which routes back out through the Daily WebRTC transport layer.

2. Optimizing Voice Activity Detection (VAD)

A critical challenge in voice AI is handling interruptions. If a user interrupts the AI while it is speaking, the agent must immediately stop generating and clear its buffer. Pipecat handles this natively via its frame management system. By tuning the VAD parameters, we ensure the agent stops talking within 100ms of detecting human speech, preventing frustrating overlapping dialogues.

Dockerizing the Infrastructure for Production

Deploying a real-time WebRTC media agent requires careful network mapping. Since WebRTC relies heavily on dynamic UDP ports for media streaming, our Docker configuration must be highly optimized.

The Dockerfile

We use a lightweight, Python-based base image optimized for asynchronous I/O operations:

FROM python:3.11-slim

WORKDIR /app

RUN apt-get update && apt-get install -y \
    build-essential \
    libsndfile1 \
    && rm -rf /var/lib/apt/lists/*

COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt

COPY . .

EXPOSE 7860

CMD ["python", "agent.py"]

The docker-compose.yml Configuration

To run the application seamlessly alongside environment variables and proper network routing, we define a docker-compose.yml file:

version: '3.8'
services:
  voice_agent:
    build: .
    env_file:
      - .env
    network_mode: "host"
    restart: always

Note: Using network_mode: "host" is highly recommended for WebRTC applications running on Linux servers. It eliminates the overhead of Docker’s internal NAT routing, shaving off crucial milliseconds from network traversal and simplifying the handling of dynamic UDP ports required by Daily Transport.

Crucial Performance Tuning for Sub-300ms Latency

Deploying the code is only half the battle. To guarantee consistency under the 300ms threshold, implement these three advanced configurations:

  • WebSocket Pre-warming: Ensure that connection handshakes with your STT and TTS providers are established the moment the user enters a session room, rather than waiting for the user to start speaking.
  • Time-to-First-Token (TTFT) Prioritization: When selecting an LLM provider, prioritize TTFT over total throughput. A model that starts streaming its first token in 80ms but has lower overall words-per-minute is vastly superior for real-time voice than a model that pauses for 400ms and then outputs text at warp speed.
  • Geographic Co-location: Deploy your Docker containers in regions physically close to your LLM/TTS provider endpoints and your target audience. If your LLM provider operates primarily out of US-East, hosting your Docker container in Singapore will add roughly 180ms of unavoidable speed-of-light network latency.

Conclusion: The Future of Conversational Interfaces

By combining the orchestration power of Pipecat, the robust WebRTC media routing of Daily Transport, and the portable efficiency of Docker, developers can comfortably break the 300ms latency barrier. The result is a conversational agent that feels naturally responsive, capable of handling real-world interruptions, dynamic pacing, and fluid dialogue.

As voice models continue to evolve from cascading pipelines into native multimodal end-to-end architectures, mastering the infrastructure required to route, containerize, and optimize these real-time streams will remain a vital competitive advantage for enterprise AI applications.

Building a Real-Time AI Voice Agent with Under 300ms Latency Using Pipecat, Daily Transport, and Docker | DPTCloud