Back to articles
Technology Insight

Building a Multi-Language AI Voice Agent for Automated Telesales: Integrating Vapi and Asterisk on a VPS

May 30, 2026

Introduction: The Dawn of AI-Driven Customer Interactions

The landscape of outbound telesales and customer support is undergoing a massive paradigm shift. Traditional automated calling systems—often reliant on rigid Interactive Voice Response (IVR) menus or pre-recorded scripts—are rapidly failing to meet the expectations of modern consumers. Today's businesses require dynamic, intelligent, and highly scalable communication channels. Enter the AI Voice Agent: an autonomous system capable of conducting fluid, multi-language conversations in real-time, understanding context, and executing business logic on the fly.

By leveraging advanced Large Language Models (LLMs) and optimized Text-to-Speech (TTS) pipelines, companies can now deploy virtual agents that sound remarkably human. However, bridging the gap between cutting-edge AI platforms and the traditional global telephony network remains a core engineering challenge. In this comprehensive guide, we will explore how to architecture and deploy a self-hosted AI Voice Agent workstation by integrating Vapi (a state-of-the-art AI voice platform) with Asterisk (the industry-standard open-source PBX framework) on a high-performance Virtual Private Server (VPS). This setup empowers your business to automate high-volume telesales campaigns globally while maintaining strict control over data routing and infrastructure costs.

The Core Architectural Components

Before diving into the configuration steps, it is essential to understand how the technical layers interact to deliver a low-latency voice experience. The entire pipeline relies on three foundational pillars:

  • Virtual Private Server (VPS): Acts as the secure, centralized hosting environment for your telephony infrastructure, ensuring high availability and dedicated network bandwidth.
  • Asterisk PBX: Handles core telecommunication routing, Session Initiation Protocol (SIP) trunking, and connection management to external telecom operators.
  • Vapi (Voice AI Layer): Manages the orchestration of Voice Activity Detection (VAD), Automatic Speech Recognition (ASR), LLM processing, and neural TTS streaming back to the caller.
Note on Latency: In conversational AI, a response delay greater than 1.5 seconds can break the illusion of a natural conversation. This architecture utilizes WebSockets and real-time SIP streaming to minimize processing overhead.

Step 1: Preparing and Securing Your VPS

To ensure crystal-clear audio quality without jitter or packet loss, your VPS must be properly provisioned. We recommend a clean installation of Ubuntu 22.04 LTS or Debian 12 on a server with at least 2 vCPUs and 4GB of RAM, preferably located near your target audience or your SIP provider's data center.

First, execute standard system updates and configure the firewall to allow SIP signaling and audio traffic:

sudo apt update && sudo apt upgrade -y
sudo ufw allow 22/tcp
sudo ufw allow 5060/udp
sudo ufw allow 10000:20000/udp
sudo ufw enable

Port 5060 is reserved for standard SIP signaling, while the 10000:20000 range accommodates the Real-time Transport Protocol (RTP) audio streams.

Step 2: Installing and Configuring Asterisk PBX

Asterisk serves as the gateway between the digital AI world and the traditional public switched telephone network (PSTN). Install Asterisk directly from the official package manager or compile it from source to ensure the latest LTS version is deployed.

sudo apt install asterisk -y

Once installed, we must configure Asterisk to communicate with your chosen SIP Trunk provider (e.g., Twilio, Telnyx, or localized telecom operators) and to route inbound/outbound traffic toward the Vapi API endpoint. This is achieved by modifying the /etc/asterisk/pjsip.conf file to define endpoints, auth configurations, and transport mechanisms. Next, we update the dialplan in /etc/asterisk/extensions.conf to instruct Asterisk on how to handle calls designated for the AI Voice Agent.

Sample Dialplan Configuration

In your extensions.conf, you will define a context that intercepts outbound telesales leads or incoming inquiries, forwarding the audio stream to Vapi via a SIP URI or custom Webhook trigger:

[telesales-ai-context]
exten => _X.,1,NoOp(Forwarding Call to Vapi AI Voice Agent)
   same => n,Answer()
   same => n,Dial(PJSIP/vapi-trunk/sip:sip-endpoint.vapi.ai)
   same => n,Hangup()

Step 3: Integrating Vapi as the AI Orchestration Layer

Vapi simplifies the complex task of stitching together speech-to-text, cognitive modeling, and text-to-speech technologies. Instead of manually writing code to connect Whisper ASR to OpenAI GPT-4, and subsequently to ElevenLabs TTS, Vapi orchestrates this flow with exceptionally low latency via a unified interface.

To connect your self-hosted Asterisk server with Vapi, log into your Vapi dashboard and perform the following actions:

  1. Navigate to the Phone Numbers section and select "Buy/Import Number" or "Create Custom SIP Trunk".
  2. Input your VPS's static IP address and the designated SIP port configuration.
  3. Configure the cryptographic security keys to ensure that only authorized SIP invitations from your Asterisk server are processed by your AI engine.

Step 4: Configuring the Multi-Language AI Agent Profile

One of the primary business advantages of this architecture is its native multi-language capability. Within the Vapi console, or via their REST API, you can define specific system prompts and choose localized neural voices tailored to your target demographics.

For a highly professional B2B telesales campaign, your agent configuration payload should explicitly outline its operational boundaries, linguistic preferences, and communication objectives:

{
  "transcriber": {
    "provider": "deepgram",
    "model": "nova-2",
    "language": "en"
  },
  "model": {
    "provider": "openai",
    "model": "gpt-4o",
    "messages": [
      {
        "role": "system",
        "content": "You are a polite, professional corporate sales representative. Your objective is to introduce our enterprise cloud solutions, handle objections with empirical data, and qualify leads for a human follow-up call."
      }
    ]
  },
  "voice": {
    "provider": "elevenlabs",
    "voiceId": "adam"
  }
}

For global deployment, you can dynamically modify the language string (e.g., to vi for Vietnamese, ja for Japanese, or es for Spanish) and swap the underlying neural voice profiles instantly based on the geographical routing logic determined by your Asterisk dialplan.

Step 5: Testing, Monitoring, and Enterprise Scalability

Before launching a large-scale outbound telesales campaign, rigorous testing is mandatory. Utilize built-in Asterisk debugging tools to monitor network traffic and analyze signal paths:asterisk -rvvv pjsip set logger on

Verify that the RTP packets are flowing correctly between your VPS and Vapi, and that audio volume levels remain consistent throughout the interaction. Once verified, you can easily scale this infrastructure. Because Asterisk handles concurrency efficiently, a standard 4-core cloud instance can easily process dozens of simultaneous, high-fidelity AI-driven telephonic conversations without performance degradation.

Conclusion: Embracing the Future of Telephony

Combining the open-source routing power of Asterisk with the advanced cognitive abilities of Vapi provides modern businesses with an unprecedented tool for automated communication. By self-hosting the core telephony engine on your own VPS, you retain full ownership of customer data paths, eliminate unnecessary middleware markups, and establish an incredibly scalable foundation for localized or international telesales campaigns. The era of repetitive, manual outbound dialing is drawing to a close—replaced by intelligent, articulate, and tireless AI voice agents.

Building a Multi-Language AI Voice Agent for Automated Telesales: Integrating Vapi and Asterisk on a VPS | DPTCloud