Back to articles
Technology Insight

Building an Automated AI Subtitle Generator: Background Whisper Translation and Hardsubbing on Linux VPS

May 30, 2026

Introduction to Automated Media Localization

In the digital age, video content is a primary driver of global engagement. However, reaching an international audience requires efficient, accurate, and scalable localization. Traditional manual subtitling and translation workflows are often slow, costly, and prone to human error. By leveraging artificial intelligence and cloud infrastructure, businesses can fully automate this pipeline.

This technical guide details how to architect and deploy a production-ready, automated AI Subtitle Generator on a Linux Virtual Private Server (VPS). Utilizing OpenAI's open-source Whisper model for automatic speech recognition (ASR) and FFmpeg for high-performance multimedia handling, this system processes video files asynchronously in the background, automatically translating and hardsubbing them without manual intervention.

---

System Architecture and Workflow Overview

Before diving into implementation, it is crucial to understand how the components interact. A robust automated system separates user interaction or file uploads from the resource-intensive processing pipeline. Below is the standard architecture for a background-processing subtitle system:

  • File Ingestion: Videos are uploaded via an API, SFTP, or a watch folder system.
  • Queue Management: A background worker monitor (such as a Bash daemon, Systemd service, or Celery task queue) detects new files and triggers processing asynchronously.
  • AI Transcription & Translation: OpenAI's Whisper processes the audio track, detects the source language, transcribes it, and optionally translates it into English or another target language, outputting SubRip (.srt) or WebVTT (.vtt) files.
  • Video Burning (Hardsubbing): FFmpeg takes the original video and overlays the generated subtitle file directly onto the video frames.
  • Notification & Delivery: The finalized video is moved to an output directory, and an automated webhook or notification is dispatched.
---

Setting Up the Linux VPS Environment

To handle deep learning inference and high-resolution video encoding, your Linux VPS should meet specific hardware guidelines. While a CPU-only server works, a VPS equipped with an NVIDIA GPU drastically accelerates Whisper's processing speeds.

Step 1: System Updates and Dependencies

First, connect to your Linux server (Ubuntu 22.04 LTS recommended) via SSH and update the package repositories. Install the core multimedia libraries and development tools required for compilation and execution:

sudo apt update && sudo apt upgrade -y
sudo apt install -y ffmpeg python3-pip python3-venv git build-essential

Verify your FFmpeg installation by running ffmpeg -version. Ensure it includes configuration flags for fonts (like fontconfig and libass), which are vital for rendering styled subtitles.

Step 2: Python Virtual Environment Setup

To avoid conflicts with system-level packages, create an isolated Python virtual environment dedicated to the AI subtitle engine:

python3 -m venv /opt/subtitle_env
source /opt/subtitle_env/bin/activate
pip install --upgrade pip
---

Installing and Configuring OpenAI Whisper

OpenAI's Whisper is available as a Python library. You can install the official repository directly using pip inside your activated virtual environment:

pip install git+[https://github.com/openai/whisper.git](https://github.com/openai/whisper.git)

Choosing the Right Whisper Model

Whisper offers five distinct model sizes, balancing speed against transcription accuracy. Selecting the right model depends heavily on your server's hardware constraints:

Model SizeParametersVRAM RequiredRelative SpeedBest Used For
tiny39 M~1 GB~32xQuick testing, low-resource environments
base74 M~1 GB~16xStandard English dictation, fast drafts
small244 M~2 GB~6xGood balance of speed and accuracy for common languages
medium769 M~5 GB~2xHigh-accuracy multilingual needs
large1550 M~10 GB1x (Base)Enterprise-grade accuracy, complex accents, and translation
Pro-tip: For automated workflows on commercial servers processing multiple languages, the medium or large-v3 model is highly recommended for optimal translation quality.
---

Building the Automated Background Processing Script

To make the system execute seamlessly in the background (headless operation), we write a master automation script. This script monitors an input directory, extracts the audio, triggers Whisper translation, and runs FFmpeg to burn the subtitles.

Create a Python script named process_video.py:import os import subprocess import whisper import sys def generate_subtitles(video_path, output_dir): print(f"[1/3] Extracting and transcribing audio from: {video_path}") model = whisper.load_model("small") # Change to medium/large based on hardware # Run Whisper inference with automatic translation to English result = model.transcribe(video_path, task="translate") # Generate SRT formatted data srt_path = os.path.join(output_dir, "subtitles.srt") with open(srt_path, "w", encoding="utf-8") as srt_file: for i, segment in enumerate(result['segments'], start=1): start = format_timestamp(segment['start']) end = format_timestamp(segment['end']) text = segment['text'].strip() srt_file.write(f"{i}\n{start} --> {end}\n{text}\n\n") return srt_path def format_timestamp(seconds): hrs = int(seconds // 3600) mins = int((seconds % 3600) // 60) secs = int(seconds % 60) ms = int((seconds % 1) * 1000) return f"{hrs:02d}:{mins:02d}:{secs:02d},{ms:03d}" def hardsub_video(video_path, srt_path, final_output): print(f"[2/3] Burning subtitles into video matrix via FFmpeg...") # Correctly escaping the subtitle path for FFmpeg's filter escaped_srt = srt_path.replace(":", "\\:").replace("'", "\\'") cmd = [ 'ffmpeg', '-y', '-i', video_path, '-vf', f"subtitles='{escaped_srt}'", '-c:a', 'copy', final_output ] subprocess.run(cmd, check=True) print(f"[3/3] Automation complete. Saved to: {final_output}") if __name__ == "__main__": if len(sys.argv) < 2: print("Usage: python process_video.py ") sys.exit(1) vid = sys.argv[1] out_directory = "/var/media/output" os.makedirs(out_directory, exist_ok=True) srt = generate_subtitles(vid, out_directory) final_vid = os.path.join(out_directory, "hardsubbed_" + os.path.basename(vid)) hardsub_video(vid, srt, final_vid)---

Configuring a Systemd Service for Unattended Background Tasks

To transform this execution script into a resilient pipeline that automatically triggers, scales, or restarts on server failure, we deploy it under a Systemd daemon process wrapper. This enables headless processing that runs even when no users are logged into the VPS.

Creating the Systemd Service Unit

Create a service file at /etc/systemd/system/subtitle-worker.service:

[Unit]
Description=AI Subtitle Generator Background Worker
After=network.target

[Service]
Type=simple
User=root
WorkingDirectory=/opt/
ExecStart=/opt/subtitle_env/bin/python /opt/process_video.py /var/media/input/target.mp4
Restart=on-failure

[Install]
WantedBy=multi-user.target

Reload the system configuration daemon, enable the service to start automatically during boot sequences, and execute the background system:

sudo systemctl daemon-reload
sudo systemctl enable subtitle-worker.service
sudo systemctl start subtitle-worker.service

To monitor real-time system performance and review detailed processing telemetry logs from Whisper and FFmpeg, use the system journal engine:

sudo journalctl -u subtitle-worker.service -f
---

Optimizing the Production Pipeline

When running machine learning pipelines at scale on a Linux server, performance bottlenecks can arise. Consider the following optimizations to maximize production efficiency:

  1. Utilize Whisper.cpp for CPU-Only Environments: If your VPS lacks an expensive dedicated GPU, rewrite the inference execution using whisper.cpp, a high-performance C/C++ port optimized specifically for x86 and ARM CPUs.
  2. FFmpeg Hardware Acceleration: Instead of relying on software encoders (libx264), utilize NVIDIA NVENC (-c:v h264_nvenc) or Intel Quick Sync to accelerate subtitle rendering speeds up to 5x.
  3. Dynamic Directory Watchers: Pair the processing script with Linux's native inotify-tools subsystem to establish dynamic folder monitoring, instantly triggering subtitle creation the second a new video file lands in the ingestion folder.

Conclusion

Building an automated AI Subtitle Generator on a Linux VPS eliminates repetitive manual workflows, standardizes digital asset localization, and vastly accelerates global media delivery timelines. By synthesizing OpenAI's powerful language transcription with the industrial processing capabilities of FFmpeg, your enterprise can deploy a scalable media localization infrastructure completely independent of expensive third-party software subscriptions.

Building an Automated AI Subtitle Generator: Background Whisper Translation and Hardsubbing on Linux VPS | DPTCloud