Back to articles
Technology Insight

Going Serverless-Like on a VPS: How to Manage Millions of Workflows with Hatchet

May 30, 2026

Introduction: The Developer's Distributed System Dilemma

Modern application development demands robust background processing. Whether you are handling heavy AI inference pipelines, processing massive batches of financial transactions, or managing complex video transcoding jobs, standard synchronous request-response cycles simply fall short. Historically, developers facing these scale requirements turned to serverless platforms like AWS Step Functions or AWS Lambda. However, as execution volume scales into millions of workflows, organizations quickly encounter a painful reality: prohibitive costs, cold starts, vendor lock-in, and restrictive timeout limits.

But what if you could achieve the elasticity, concurrency control, and developer experience of serverless workflows directly on your own cost-effective Virtual Private Server (VPS)? Enter Hatchet. Hatchet is an open-source, high-performance distributed workflow engine designed to replace traditional message queues and rigid orchestrators. This guide provides a comprehensive roadmap to building a 'Serverless-like' architecture on a standard VPS, empowering you to manage millions of concurrent workflows efficiently and predictably.

The Core Challenge of Background Processing at Scale

Before diving into the technical implementation, it is vital to understand why traditional task queues like Celery, BullMQ, or Sidekiq struggle under multi-million workflow demands:

  • Lack of Fine-Grained Concurrency Control: Standard queues often operate on a simple first-in, first-out (FIFO) basis or require complex redis architectures to enforce per-user or per-tenant rate limits.
  • Poor Visibility and Observability: Debugging a nested, multi-step asynchronous workflow in traditional queues usually requires digging through disjointed application logs or setting up complex APM tools.
  • Resource Inefficiency: Maintaining idle workers to handle sudden, unpredictable spikes in traffic wastes valuable VPS CPU and memory resources.

Hatchet addresses these exact pain points by decoupling the orchestration engine from the execution workers. Written in Go and powered by a high-throughput PostgreSQL engine, Hatchet can process thousands of events per second with sub-millisecond overhead, maximizing the utility of every gigabyte of RAM on your VPS.

Architecting the 'Serverless-Like' Model on a VPS

In a standard serverless setup, cloud providers provision micro-virtual machines on demand. In our Serverless-like VPS architecture, we replicate this efficiency by separating the control plane from the data plane. The Hatchet Engine acts as the centralized control plane, while your lightweight application services act as dynamic, long-lived workers that scale internally via multi-threading or green threads.

Key Architectural Components:

  1. The Hatchet Engine: A centralized binary running on your VPS (or via Docker Compose) that manages workflow states, cron schedules, retries, and event distributions.
  2. The Database: A highly optimized PostgreSQL instance that stores workflow graphs, step states, and execution histories.
  3. The Workers: Your application code (Node.js, Python, or Go) running on the same VPS or distributed across multiple nodes. These workers connect to the Hatchet Engine via secure, persistent gRPC streams.
Note on Efficiency: Because Hatchet uses gRPC for bi-directional streaming, communication overhead between the orchestrator and your workers is virtually non-existent, ensuring low latency and immediate task allocation.

Step-by-Step Implementation Guide

Step 1: Deploying the Hatchet Infrastructure

The fastest and most reliable way to get Hatchet running on a standard Linux VPS is through Docker Compose. This ensures isolated, reproducible environments for both the engine and its dependencies.

version: '3.8'
services:
  postgres:
    image: postgres:16-alpine
    environment:
      POSTGRES_USER: hatchet
      POSTGRES_PASSWORD: your_secure_password
      POSTGRES_DB: hatchet
    volumes:
      - pgdata:/var/lib/postgresql/data
    ports:
      - "5432:5432"

  hatchet-engine:
    image: ghcr.io/hatchet-dev/hatchet/hatchet-engine:latest
    environment:
      - HATCHET_CLIENT_POSTGRES_URL=postgresql://hatchet:your_secure_password@postgres:5432/hatchet?sslmode=disable
    ports:
      - "8080:8080"
      - "7070:7070"
    depends_on:
      - postgres

volumes:
  pgdata:

Run docker compose up -d to initialize the backend services. Once active, the engine exposes a web dashboard on port 8080 and the gRPC communication gateway on port 7070.

Step 2: Defining a Robust, Multi-Step Workflow

Hatchet allows developers to define complex workflows directly in code using familiar paradigms. Below is a practical example of a Python-based asynchronous user onboarding workflow, featuring dependent steps, custom retries, and timeouts.

from hatchet_sdk import Hatchet

hatchet = Hatchet()

@hatchet.workflow(on_events=["user:registered"])
class UserOnboardingWorkflow:
    
    @hatchet.step(timeout="30s", retries=3)
    def provision_account(self, context):
        input_data = context.workflow_input()
        user_id = input_data.get("user_id")
        # Simulate internal API call
        return {"status": "account_created", "uid": user_id}

    @hatchet.step(parents=["provision_account"])
    def send_welcome_email(self, context):
        step_result = context.step_output("provision_account")
        # Execute email logic
        return {"email_sent": True}

In this snippet, the send_welcome_email step will only execute after provision_account resolves successfully. If the account provisioning fails due to an external network glitch, Hatchet automatically handles the three configured retries exponentially before marking the workflow as failed.

Strategies for Scaling to Millions of Executions

Running millions of jobs on a single VPS requires deliberate resource management. To optimize performance and ensure high availability, leverage Hatchet's advanced concurrency features:

1. Fine-Grained Concurrency Limits

To prevent external third-party APIs from blocking or rate-limiting your worker processes, you can implement strict concurrency limits at the workflow or tenant level. For instance, you can limit OpenAI API calls to a maximum of 50 concurrent executions across your entire system, buffering any excess requests in Hatchet's internal queue without consuming extra RAM.

2. Leveraging Ephemeral Workers

Because Hatchet workers communicate via gRPC, they do not need to be co-located with the main engine database. If your VPS experiences a heavy processing surge, you can spin up temporary, lightweight compute nodes on cheaper cloud micro-instances, point them to your main VPS Hatchet gRPC endpoint, and scale your processing power instantly.

3. Database Optimization and Maintenance

When executing millions of workflows, your PostgreSQL instance will expand rapidly. It is highly recommended to implement database partitioning on Hatchet's execution log tables and regularly run background vacuuming processes to keep disk read/write operations optimal.

Conclusion: Embracing High-Performance Sovereignty

Building a 'Serverless-like' infrastructure with Hatchet on a VPS offers the ultimate balance between cloud flexibility and bare-metal cost predictability. By taking control of your orchestration layer, you eliminate arbitrary timeout limits, bypass expensive cloud vendor premiums, and retain full sovereignty over your application data.

Whether you are managing a few thousand complex automation sequences or scaling up to hundreds of millions of background events, Hatchet provides the modern foundation needed to run dependable distributed systems without breaking the bank.

Going Serverless-Like on a VPS: How to Manage Millions of Workflows with Hatchet | DPTCloud