Back to articles
Technology Insight

Building a Self-Hosted AI Voice-Agent Call Center on a VPS: Integrating Asterisk, Whisper Live, and Llama-3-Matcha for Automated Sales

May 26, 2026

Introduction: The Shift Toward Private, Autonomous Voice Automation

In the competitive landscape of modern telecommunications and customer acquisition, the call center remains a vital touchpoint. However, traditional call centers face steep challenges: high agent turnover, escalating operational costs, and human limitations in handling massive outbound volumes simultaneously. While cloud-based AI voice solutions exist, they often come with prohibitive per-minute pricing and raise severe data privacy concerns regarding proprietary customer data.

The alternative? A fully self-hosted, enterprise-grade AI Voice-Agent Call Center deployed on a Virtual Private Server (VPS). By converging open-source telephony with cutting-edge open-weights Large Language Models (LLMs), businesses can deploy autonomous voice agents capable of qualifying leads, answering complex product queries, and closing orders in real-time. This comprehensive technical guide details how to architecture and build this pipeline using Asterisk PBX, Whisper Live, and Llama-3-Matcha.

The Core Architectural Blueprint

To achieve a seamless, low-latency conversation that mimics human interaction, the system relies on an optimized three-tier architecture. High latency is the ultimate killer of voice AI user experience; therefore, data must flow continuously through a streaming pipeline rather than waiting for structural silences.

  • Telephony Layer (Asterisk PBX): Manages inbound and outbound SIP trunks, handles call routing, and streams raw audio bidirectionally via WebSockets or AudioSocket.
  • Transcription Layer (Whisper Live): Consumes the live audio stream from Asterisk, utilizing optimized Whisper weights to output text tokens with sub-100ms latency.
  • Cognitive & Synthesis Layer (Llama-3-Matcha): An ultra-fast, fine-tuned LLM that processes the text context, determines the conversational next step, and drives a specialized Text-to-Speech (TTS) engine to stream audio back to the caller.
Data Privacy Advantage: By hosting this stack on your own infrastructure, zero customer voice prints or conversational logs are transmitted to external third-party AI vendors, ensuring strict compliance with GDPR and local data protection regulations.

Step 1: Setting Up the Telephony Foundation with Asterisk

Asterisk serves as the robust backbone of our call center infrastructure. To stream audio in real-time to our AI model, we configure Asterisk to utilize its built-in AudioSocket or an external WebSocket gateway. This bypasses traditional file-based recording storage and feeds audio directly into memory streams.

First, configure your SIP trunk provider in pjsip.conf to receive and make external calls. Next, establish the routing logic within the dialplan (extensions.conf). When an outbound call connects, Asterisk triggers an application extension that initiates a bidirectional audio stream directed toward the local IP and port of the Whisper Live daemon. This layout ensures that as soon as the customer says "Hello," the raw PCM audio bytes are instantly captured.

Step 2: Real-Time Transcription via Whisper Live

Standard implementations of OpenAI's Whisper require an entire audio file to be completed before processing. For a live conversation, this is unacceptable. Whisper Live overcomes this barrier by utilizing a rolling voice-activity detection (VAD) window and stateful audio chunking to transcribe speech on the fly.

Deployed inside a Docker container on your VPS, Whisper Live listens to the incoming stream from Asterisk. It constantly processes incoming 16kHz audio chunks. To achieve maximum throughput on standard VPS hardware, we implement TensorRT-LLM or ggml optimizations to run quantized variants of the Whisper model (such as Whisper-Medium or Whisper-Small fine-tuned for the local language). The resulting text strings are continuously emitted via a localized WebSocket server directly to the orchestration engine.

Step 3: Intelligence Orchestration with Llama-3-Matcha

Once the text is generated, it is passed to the brain of our autonomous agent: Llama-3-Matcha. This specialized iteration of Meta's Llama-3 framework is highly optimized for fast inference and transactional dialogues, making it exceptionally well-suited for automated sales scripts and order placement processes.

Prompt Engineering for Sales Conversions

To ensure the model successfully steers conversations toward an order confirmation, it must be constrained by strict system prompts. The prompt defines the agent's persona, product boundaries, and specific API triggers. For instance, when a customer confirms an item selection, Llama-3-Matcha recognizes the intent and structured format, triggering a backend webhook to verify inventory or lock in a CRM order ID before responding vocally.

Mitigating Latency through Token Streaming

To maintain conversational fluidity, Llama-3-Matcha does not wait for a complete sentence to generate. It utilizes streaming generation tokens. As the first few words of the response are computed, they are sent immediately to an optimized Text-to-Speech (TTS) framework (such as StyleTTS2 or a fast vits engine), which converts text to audio chunks and passes them right back down the Asterisk channel to the listener's ear.

Deployment and Optimization Strategies on VPS

Running an end-to-end voice AI pipeline requires careful resource allocation on your VPS provider. To ensure continuous service without call drops or robotic audio distortion, implement the following deployment principles:

  1. Compute Allocation: Allocate dedicated GPU slices (e.g., NVIDIA A10G or T4 instances) if handling high concurrency. For purely CPU-bound VPS infrastructure, utilize highly quantized models (4-bit or 5-bit GGUF formats) and limit simultaneous call channels.
  2. Process Isolation: Use Docker Compose to isolate Asterisk, Whisper Live, and the LLM inference server into distinct containers. Apply strict CPU pinning to Asterisk to prevent conversational AI spikes from degrading audio package delivery.
  3. Caching & State Management: Implement a Redis layer to manage live session states, customer profiles, and dialogue history. This prevents the LLM from losing context during longer calls.

Conclusion: The Future of Autonomous Customer Engagement

Building a self-hosted AI Voice-Agent Call Center marks a paradigm shift in how businesses handle high-volume outbound and inbound voice operations. By integrating the rock-solid routing of Asterisk, the instant transcription capabilities of Whisper Live, and the contextual intelligence of Llama-3-Matcha, you establish an automated sales asset that operates 24/7 with zero variable API costs. As open-source voice models continue to evolve, owning your infrastructure guarantees that your enterprise remains agile, secure, and uniquely positioned to scale customer acquisition seamlessly.

Building a Self-Hosted AI Voice-Agent Call Center on a VPS: Integrating Asterisk, Whisper Live, and Llama-3-Matcha for Automated Sales | DPTCloud