Back to articles
Technology Insight

Building an Automated AI Voice Agent Cold Calling Server with Asterisk and Llama-3-Matcha on a VPS

May 26, 2026

Introduction: The Evolution of Outbound Sales Infrastructure

In the highly competitive landscape of modern business, outbound lead generation remains a critical driver of revenue. However, traditional cold calling is fraught with inefficiencies, high agent turnover, and skyrocketing operational costs. The convergence of open-source telecommunications and advanced Generative AI has birthed a paradigm shift: Automated AI Voice Agents. By replacing or augmenting human dialers with intelligent, context-aware AI agents, enterprises can scale their outreach infinitely while maintaining a consistent, professional brand voice.

This technical guide provides a comprehensive blueprint for building a production-grade AI Voice Agent Cold Calling Server. By leveraging the robust telephony capabilities of Asterisk PBX and the specialized, low-latency conversational power of the Llama-3-Matcha Large Language Model (LLM), we will construct an automated pipeline capable of dialing prospects, detecting human responses, engaging in fluid multi-turn dialogue, and logging qualification data—all hosted cost-effectively on a standard Virtual Private Server (VPS).

System Architecture Overview

Building a real-time voice AI system requires careful orchestrations to minimize latency. High latency is the ultimate killer of conversational AI; a delay of more than 1.5 seconds between a user speaking and the AI responding breaks the illusion of natural conversation and leads to immediate hangups. Our architecture is designed to optimize data flow across four core layers:

  • Telephony Layer (Asterisk): Handles Session Initiation Protocol (SIP) trunking, inbound/outbound call routing, and media streaming.
  • Audio Processing Pipeline (Speech-to-Text & Text-to-Speech): Converts the incoming real-time audio stream (PCM) into text via ultra-fast Whisper models (like Whisper-Live or Whisper-Streaming) and converts the AI's textual response back to fluid speech via tools like Coqui TTS or Kokoro.
  • Orchestration Layer (Python/Node.js Node-ARI): Acts as the central nervous system, interfacing with Asterisk via the Asterisk REST Interface (ARI) to capture audio events and coordinate with the AI models.
  • Cognitive Layer (Llama-3-Matcha): The brain of the agent, responsible for understanding context, adhering to sales scripts, and generating concise responses optimized for voice interaction.
Why Llama-3-Matcha? While standard LLMs are trained primarily for text generation, 'Matcha' variants and finely tuned small language models (SLMs) are optimized for speed, low token-to-first-token (TTFT) latency, and concise conversational turns—making them ideal for voice applications running on limited VPS hardware.

Step 1: Setting Up the VPS and Installing Asterisk

To ensure smooth operations, your VPS should feature at least 4 vCPUs, 8GB of RAM, and an Ubuntu 22.04 LTS or 24.04 LTS environment. If you plan to run local LLM inference, a GPU-accelerated VPS (such as those offering NVIDIA T4 or L4 instances) is highly recommended. However, for cost optimization, this guide covers an architecture where the VPS runs the lightweight components locally and interfaces with an optimized inference endpoint (such as vLLM or Ollama) for the language model.

Installing Asterisk PBX

First, update your package repository and install Asterisk along with its essential dependencies:

sudo apt update && sudo apt upgrade -y
sudo apt install asterisk asterisk-dev -y

Once installed, enable and start the Asterisk service to ensure it runs automatically on system boot:

sudo systemctl enable asterisk
sudo systemctl start asterisk

To verify the installation, enter the Asterisk Command Line Interface (CLI):

asterisk -rvvv

Step 2: Configuring Asterisk REST Interface (ARI) and SIP Trunking

Traditional dialplans (extensions.conf) are too rigid for dynamic AI routing. Instead, we utilize the Asterisk REST Interface (ARI), which allows our external Python or Node.js orchestration script to take absolute control of a call channel, intercept media streams, and inject audio dynamically.

Enabling ARI

Edit the /etc/asterisk/ari.conf file to enable the HTTP server and define secure credentials for your orchestration application:

[general]
enabled = yes
pretty = yes

[ai_agent_user]
password = YourSecurePasswordHere
read_only = no

Next, ensure the internal HTTP server is active by editing /etc/asterisk/http.conf:

[general]
enabled = yes
bindaddr = 0.0.0.0
bindport = 8088

Configuring the Outbound SIP Trunk

To make external phone calls to prospects, you must connect Asterisk to a wholesale VoIP provider via PJSIP. Configure your outbound endpoint in /etc/asterisk/pjsip.conf:

[transport-udp]
type=transport
protocol=udp
bind=0.0.0.0

[sip_provider]
type=registration
outbound_auth=sip_provider_auth
server_uri=sip:sip.yourprovider.com
client_uri=sip:[email protected]

[sip_provider_auth]
type=auth
auth_type=userpass
username=your_username
password=your_sip_password

[sip_provider_endpoint]
type=endpoint
context=from-external
disallow=all
allow=ulaw,alaw
aors=sip_provider
auth=sip_provider_auth

Step 3: Deploying Llama-3-Matcha for Low-Latency Conversational Turn-Taking

The core conversational engine relies on Llama-3-Matcha, an optimized model trained to handle colloquial, verbal communication rather than overly formal, essays-style text. To maintain a conversational flow, we host this model using vLLM or Ollama to utilize token streaming, allowing us to feed text to our Text-to-Speech engine before the model finishes generating the entire sentence.

Running Llama-3-Matcha via Ollama

Install Ollama on your server or an isolated GPU instance:

curl -fsSL [https://ollama.com/install.sh](https://ollama.com/install.sh) | sh

Pull the optimized model weights:

ollama run llama3-matcha

To fit a cold-calling script, we must inject a precise System Prompt that dictates constraints: brevity, persuasion, structured objection handling, and immediate question-asking to retain control of the call. Here is an example of a system prompt structure:

"You are an automated B2B sales development representative for Acme Cloud Solutions. Your goal is to qualify if the prospect faces server downtime issues and schedule a 10-minute demo. Keep responses under 20 words. Never use bullet points. Sound casual, empathetic, and professional. Wait for the user to finish speaking before responding."

Step 4: Building the Orchestration Pipeline (The Python 'Bridge')

With Asterisk ready to initiate calls and Llama-3-Matcha ready to think, we build the orchestration layer using Python. This script connects to Asterisk's ARI via WebSockets, initiates an outbound call via the SIP trunk, and hooks into the media stream using an external media application protocol (like Audiosocket or RTP streaming).

The Technical Flow of a Call

  1. Initiation: The Python script triggers an ARI command to dial a prospect's number via the SIP provider.
  2. Answer Detection: Once the channel state changes to "Up" (answered), Asterisk transfers the call into a Stasis application bridge.
  3. Audio Ingestion: The script reads raw PCM audio from the Asterisk channel and streams it directly to a fast Speech-to-Text engine (e.g., Whisper-Live).
  4. VAD (Voice Activity Detection): Silero VAD is used to detect when the user stops speaking. This acts as the trigger to send the accumulated transcript to Llama-3-Matcha.
  5. Inference & Synthesis: Llama-3-Matcha streams text tokens to the TTS engine. The TTS engine synthesizes audio chunks on the fly and writes them back to the Asterisk channel playback queue.

Optimization and Production Considerations

Transitioning from a proof-of-concept to a production-scale automated dialer requires addressing several real-world operational challenges:

Answering Machine Detection (AMD)

Calling prospects results in hitting voicemails at least 60-70% of the time. Running an LLM against a pre-recorded voicemail greeting wastes computational resources and incurs unnecessary SIP costs. Configure Asterisk's native app_amd module to analyze initial audio patterns. If a machine greeting is detected, the script should hang up immediately or drop a pre-recorded voicemail message without engaging the cognitive layer.

Handling Latency Bottlenecks

To keep the total latency under 1.2 seconds, implement the following optimizations:

  • Audio Chunk Size: Use 20ms or 30ms audio frame sizes when passing data to the VAD and STT modules to avoid processing delays.
  • Model Quantization: Run Llama-3-Matcha in a quantized format (e.g., INT4 or INT8) to dramatically boost generation speed without sacrificing context comprehension.
  • Sentence Splitting for TTS: Do not wait for the entire LLM response. Split output strings by punctuation markers (periods, commas, question marks) and pass individual sentences to the TTS pipeline sequentially.

Conclusion

Building a custom, automated AI Voice Agent Server using Asterisk and Llama-3-Matcha empowers businesses to take full control over their automated outbound sales infrastructure. By removing reliance on expensive, closed-source third-party Voice AI APIs, organizations can dramatically slash per-minute operational costs, ensure absolute data privacy, and fine-tune their proprietary models to match unique industry verticals. With open-source telecommunications and modern LLMs working in unison, the future of highly efficient, human-like automated scaling is within your immediate grasp.

Building an Automated AI Voice Agent Cold Calling Server with Asterisk and Llama-3-Matcha on a VPS | DPTCloud