Back to articles
Technology Insight

Building an AI-Driven Automated Invoice Processing System on VPS using Donut Transformer

May 26, 2026

Introduction: The Evolution of Invoice Processing

In the modern business landscape, efficiency and accuracy are the twin pillars of operational excellence. For finance and accounting departments, manual data entry remains a notorious bottleneck. Traditional invoice processing relies heavily on human operators to read, interpret, and manually log data into Enterprise Resource Planning (ERP) or accounting systems. This approach is not only slow and costly but also prone to human error—mistakes that can lead to costly compliance issues or strained vendor relationships.

While optical character recognition (OCR) systems offered a first-generation solution, they frequently fail when encountering semi-structured or unstructured documents with varied layouts. Enter the era of Intelligent Document Processing (IDP). By leveraging advanced deep learning architectures, businesses can now build AI-Driven Automated Invoice Processing systems. In this technical guide, we will explore how to implement and deploy an end-to-end invoice parsing solution utilizing the groundbreaking Donut (Document Understanding Transformer) model, hosted independently on a scalable Virtual Private Server (VPS).

The Core Tech Stack: Why Donut and VPS?

Building an enterprise-grade automation pipeline requires choosing the right balance between model performance and infrastructure efficiency. Our architecture relies on two foundational components:

1. Donut Transformer Model

Unlike traditional pipeline-based IDP systems that first convert images to text via OCR and then apply Named Entity Recognition (NER), the Donut model introduces an OCR-free end-to-end approach. Developed by NAVER AI Lab, Donut utilizes a vision-to-text encoder-decoder structure based on the Swin Transformer and BART. It directly reads the visual features of an invoice image and maps them to a structured JSON output. This eliminates OCR errors caused by poor resolution, stylized fonts, or artifacts, resulting in superior accuracy and faster processing speeds.

2. Self-Hosted VPS Deployment

While cloud-based AI APIs offer convenience, they introduce significant long-term recurring costs, API rate limits, and potentially sensitive data privacy concerns. Deploying your solution on a self-hosted VPS (such as DigitalOcean, Linode, or AWS EC2) ensures absolute control over your data pipeline, guarantees strict compliance with financial privacy regulations, and provides predictable, fixed operational budgeting.

Architectural Overview of the Pipeline

An automated invoice processing system requires a structured pipeline to seamlessly transform raw files into actionable business data. The core architecture consists of the following four sequential phases:

  1. Ingestion and Preprocessing: PDFs or scanned images of invoices are uploaded via an API endpoint or monitored directory. The system normalizes the image size, DPI, and format to match the input specifications of the vision encoder.
  2. Deep Learning Inference: The preprocessed image is fed into the Donut model, where the Swin Transformer extracts visual features and the text decoder autoregressively generates a structured string.
  3. Post-processing & Validation: The raw output string is parsed into a standard JSON object. Validation rules (e.g., checking if the calculated tax matches the subtotal) are applied automatically.
  4. Downstream Integration: The validated data is pushed directly to an accounting database, ERP platform, or Webhook endpoint for immediate consumption.

Step-by-Step Implementation Guide

Let us walk through the technical steps required to prepare your VPS environment and implement the extraction core using Python and the Hugging Face Transformers library.

Step 1: Preparing your VPS Environment

To ensure optimal inference latency, your VPS should ideally be equipped with a modern GPU. However, for low-to-medium volume processing, a high-performance, multi-core CPU VPS with optimized inference runtimes can suffice. Begin by updating your system and installing the necessary system dependencies:

sudo apt-get update && sudo apt-get upgrade -y
sudo apt-get install python3-pip python3-venv libgl1-mesa-glx -y

Next, create an isolated virtual environment and install the required machine learning frameworks:

python3 -m venv venv
source venv/bin/activate
pip install --upgrade pip
pip install torch torchvision --index-url [https://download.pytorch.org/whl/cpu](https://download.pytorch.org/whl/cpu)
pip install transformers transformers[sentencepiece] Pillow pydantic
Note: If your VPS includes an NVIDIA GPU, replace the PyTorch installation command with the appropriate CUDA-enabled version to accelerate inference speeds.

Step 2: Loading the Model and Running Inference

With the environment configured, we can now write the Python script to initialize the Donut processor and model, load an invoice image, and extract structured data. We will utilize a model pre-trained or fine-tuned on document parsing tasks, such as naver-clova-ix/donut-base-finetuned-cord-v2 or a domain-specific invoice model from the Hugging Face Hub.

import torch
from PIL import Image
from transformers import DonutProcessor, VisionEncoderDecoderModel
import re
import json

# Initialize processor and model
processor = DonutProcessor.from_pretrained("naver-clova-ix/donut-base-finetuned-cord-v2")
model = VisionEncoderDecoderModel.from_pretrained("naver-clova-ix/donut-base-finetuned-cord-v2")

device = "cuda" if torch.cuda.is_available() else "cpu"
model.to(device)

def process_invoice(image_path):
    # Load and convert image
    image = Image.open(image_path).convert("RGB")
    
    # Prepare task prompt
    task_prompt = ""
    decoder_input_ids = processor.tokenizer(task_prompt, add_special_tokens=False, return_tensors="pt").input_ids
    
    # Process image
    pixel_values = processor(image, return_tensors="pt").pixel_values
    
    # Generate output
    outputs = model.generate(
        pixel_values.to(device),
        decoder_input_ids=decoder_input_ids.to(device),
        max_length=model.config.decoder.max_position_embeddings,
        pad_token_id=processor.tokenizer.pad_token_id,
        eos_token_id=processor.tokenizer.eos_token_id,
        use_cache=True,
        bad_words_ids=[[processor.tokenizer.unk_token_id]],
        return_dict_in_generate=True,
    )
    
    # Decode result
    sequence = processor.batch_decode(outputs.sequences)[0]
    sequence = sequence.replace(processor.tokenizer.eos_token, "").replace(processor.tokenizer.pad_token, "")
    sequence = re.sub(r"<.*?>", "", sequence, count=1).strip()  # remove first task start token
    
    return processor.token2json(sequence)

# Example usage
# data = process_invoice("path_to_invoice.jpg")
# print(json.dumps(data, indent=4))

Optimizing Production Deployment on VPS

Moving from a localized script to a resilient, production-grade business tool requires addressing scaling, monitoring, and stability challenges on your VPS:

  • Containerization with Docker: Wrap your application, dependencies, and model weights inside a Docker container. This ensures environment consistency between your local machine and your VPS, speeding up CI/CD deployments.
  • Asynchronous Queue Management: Deep learning inference is a resource-intensive operation. Do not process invoices synchronously inside an HTTP request. Instead, use a task manager like Celery paired with Redis or RabbitMQ. The API should immediately acknowledge receipt, place the document in a queue, and let background worker processes manage the CPU/GPU workload.
  • Model Optimization: If utilizing a CPU-only VPS, optimize inference via quantization (e.g., converting model weights from FP32 to INT8) or leverage the OpenVINO toolkit. This can reduce memory footprints by up to 60% and double inference speeds.

Conclusion: Driving ROI Through AI Native Automation

Transitioning from error-prone manual processing to an end-to-end AI-Driven Automated Invoice Processing platform on a VPS is an investment that yields compounding dividends. By removing the traditional OCR layer and leveraging the holistic vision-to-text capacity of the Donut Transformer, your organization gains superior extraction quality even on complex documents. Furthermore, maintaining this infrastructure internally on a VPS guarantees data compliance, preserves strict financial confidentiality, and fundamentally drives down operational overhead. As intelligent automation redefines modern corporate workflows, self-hosted IDP systems stand out as a foundational pillar of competitive digital strategies.

Building an AI-Driven Automated Invoice Processing System on VPS using Donut Transformer | DPTCloud