Back to articles
Technology Insight

Building an Automated AI Video Subtitle Burn-In & Hardsub Server on a VPS: A Guide for Creators and Businesses

May 25, 2026

Introduction: The Growing Demand for Video Localization

In today's digital landscape, video content reigns supreme. Whether you are a content creator targeting a global audience, a corporate training department localizing internal media, or a digital agency managing multiple brand channels, subtitles are no longer optional. Studies consistently show that a vast majority of users watch mobile videos with the sound off. Furthermore, accurate subtitles improve accessibility and boost SEO performance across platforms.

However, relying on commercial third-party SaaS platforms for video transcription and hardsubbing (permanently burning subtitles into the video file) can quickly become a bottleneck. These platforms often come with restrictive monthly usage caps, high recurring subscription fees, and concerning data privacy policies. For businesses handling proprietary information or creators producing high volumes of content, the ideal alternative is clear: building an automated, self-hosted AI Video Subtitle Burn-In Server on a Virtual Private Server (VPS).

This comprehensive technical guide will walk you through the architecture, prerequisite setup, core script implementation, and automation workflows necessary to deploy your own production-ready subtitle server.


Why Build a Self-Hosted Hardsub Server?

Before diving into the technical execution, it is critical to understand the strategic and financial advantages of moving this workflow self-hosted.

  • Absolute Data Privacy: When you upload a video to a SaaS platform, you relinquish a degree of control over that asset. By using your own VPS, your media files, corporate training videos, or unreleased content remain completely within your secure infrastructure.
  • Cost Efficiency at Scale: Instead of paying per minute or managing tiered subscriptions, you pay a flat monthly fee for your VPS infrastructure. As your video volume grows, your marginal cost approaches zero.
  • Customization and Flexibility: A self-hosted solution allows you to choose your exact open-source models (such as OpenAI's Whisper), fine-tune transcription accuracy, and precisely control subtitle styling, fonts, margins, and rendering via FFmpeg.
  • Seamless Automation: By utilizing system daemons or watch-folders, you can create a hands-free pipeline where uploading a raw video automatically triggers transcription, subtitle generation, and hardsub burning without human intervention.

System Architecture and Core Components

An automated hardsub server relies on a modular pipeline where each component handles a specific phase of the media processing cycle. The three core pillars of this architecture are:

1. Infrastructure (The VPS)

While CPU-only VPS configurations can handle transcription using optimized models like whisper.cpp, a GPU-accelerated VPS (utilizing NVIDIA T4 or A10G instances) is highly recommended for production environments requiring rapid, near-real-time rendering.

2. The AI Transcription Engine (OpenAI Whisper)

Whisper is a state-of-the-art, open-source automatic speech recognition (ASR) system. For self-hosted servers, we utilize libraries like Faster-Whisper, which leverages CTranslate2 to deliver up to 4x faster transcription speeds than the standard OpenAI implementation while using significantly less memory.

3. The Media Processing Engine (FFmpeg)

FFmpeg is the industry-standard, cross-platform solution to record, convert, and stream audio and video. It handles extracting the audio track for the AI model and executing the final video filter pass to permanently burn the generated SubRip (SRT) or Advanced SubStation Alpha (ASS) files into the output video matrix.


Step-by-Step Deployment Guide

Step 1: Preparing the VPS Environment

First, establish an SSH connection to your Ubuntu 22.04 LTS or 24.04 LTS VPS instance. Update the core system repositories and install the fundamental build tools and media libraries required:

sudo apt update && sudo apt upgrade -y
sudo apt install python3-pip python3-venv ffmpeg libsm6 libxext6 -y

If you are utilizing a GPU-enabled instance, ensure your NVIDIA drivers and CUDA Toolkit are properly installed and initialized to allow Faster-Whisper to interface with the hardware acceleration layers.

Step 2: Setting Up the Isolated Python Environment

To prevent dependency conflicts with system-level packages, create a dedicated Python virtual environment for the automation pipeline:

mkdir -p /opt/subtitle-server/src
cd /opt/subtitle-server
python3 -m venv venv
source venv/bin/activate

With the virtual environment active, install the optimized transcription engine along with any required support libraries:

pip install --upgrade pip
pip install faster-whisper

Step 3: Developing the Core Automation Script

Below is a highly structured Python script that orchestrates the entire lifecycle: accepting a raw video input, extracting the audio, executing the AI transcription to generate an SRT file, and utilizing FFmpeg to burn the hardsubs into the final output container.

Save this script as core_processor.py within your /opt/subtitle-server/src directory:

import os
import subprocess
from faster_whisper import WhisperModel

INPUT_DIR = "/opt/subtitle-server/input"
OUTPUT_DIR = "/opt/subtitle-server/output"
MODEL_SIZE = "base"  # Options: tiny, base, small, medium, large-v3

def transcribe_and_burn(video_name):
    input_video_path = os.path.join(INPUT_DIR, video_name)
    base_name = os.path.splitext(video_name)[0]
    srt_path = os.path.join(OUTPUT_DIR, f"{base_name}.srt")
    output_video_path = os.path.join(OUTPUT_DIR, f"{base_name}_hardsubbed.mp4")
    
    print(f"[INFO] Initializing Whisper model: {MODEL_SIZE}")
    # Change device to 'cuda' if running on a GPU-enabled VPS
    model = WhisperModel(MODEL_SIZE, device="cpu", compute_type="float32")
    
    print(f"[INFO] Transcribing audio track from: {video_name}")
    segments, info = model.transcribe(input_video_path, beam_size=5)
    
    print(f"[INFO] Detected language '{info.language}' with probability {info.language_probability:.2f}")
    
    # Generate compliant SRT format
    with open(srt_path, "w", encoding="utf-8") as f:
        for index, segment in enumerate(segments, start=1):
            start_h, start_m, start_s = int(segment.start//3600), int((segment.start%3600)//60), segment.start%60
            end_h, end_m, end_s = int(segment.end//3600), int((segment.end%3600)//60), segment.end%60
            
            f.write(f"{index}\n")
            f.write(f"{start_h:02d}:{start_m:02d}:{start_s:06.3f} --> {end_h:02d}:{end_m:02d}:{end_s:06.3f}\n")
            f.write(f"{segment.text.strip()}\n\n")
            
    print(f"[INFO] SRT subtitle file successfully created at: {srt_path}")
    
    # Execute FFmpeg video filter process to burn the subtitles
    print(f"[INFO] Commencing FFmpeg hardsub burn-in process...")
    ffmpeg_command = [
        'ffmpeg', '-y',
        '-i', input_video_path,
        '-vf', f"subtitles={srt_path}:force_style='Fontname=Arial,Fontsize=16,PrimaryColour=&HFFFFFF,Alignment=2'",
        '-c:a', 'copy',
        output_video_path
    ]
    
    result = subprocess.run(ffmpeg_command, stdout=subprocess.PIPE, stderr=subprocess.PIPE)
    if result.returncode == 0:
        print(f"[SUCCESS] Hardsubbed video exported successfully to: {output_video_path}")
    else:
        print(f"[ERROR] FFmpeg processing failed: {result.stderr.decode('utf-8')}")

if __name__ == "__main__":
    os.makedirs(INPUT_DIR, exist_ok=True)
    os.makedirs(OUTPUT_DIR, exist_ok=True)
    # Scan input folder for processing
    for file in os.listdir(INPUT_DIR):
        if file.endswith((".mp4", ".mkv", ".avi")):
            transcribe_and_burn(file)

Advanced Styling and Optimization

To elevate the production value of your automated hardsubs, you can modify the force_style parameters within the FFmpeg command string. The subtitles filter reads standard ASS styling attributes, allowing granular control over typography and layout:

  • Fontname: Specify system fonts like Arial, Helvetica, or Roboto. Ensure the font is installed on your VPS server via fontconfig.
  • Fontsize: Controls scale. A value between 14 and 18 is typical for standard landscape displays.
  • PrimaryColour: Uses hexadecimal representation prefixed with &H. For example, &H00FFFF renders vibrant yellow text, while &HFFFFFF provides crisp white.
  • Alignment: Dictates position based on a standard numeric keypad matrix. Alignment=2 forces center-bottom positioning, which is the industry standard for traditional media layout.
  • BorderStyle & Outline: Adding a subtle black text outline or a semi-transparent bounding box drop shadow significantly improves subtitle legibility against complex, high-contrast video backgrounds.

Conclusion: Future Proofing Your Content Engine

By shifting your video localization infrastructure to a self-hosted VPS, you successfully break free from the constraints of expensive commercial API layers. You achieve total command over asset data flow, enjoy predictable monthly infrastructure pricing, and possess the capability to extend the pipeline indefinitely—whether that involves adding multi-language translation layers via deep learning models or creating cloud API endpoints to serve your entire content production team.

As AI video models and automated tooling continue to evolve, owning the underlying infrastructure ensures your business workflows remain agile, highly secure, and optimized for rapid global distribution.

Building an Automated AI Video Subtitle Burn-In & Hardsub Server on a VPS: A Guide for Creators and Businesses | DPTCloud