Back to articles
Technology Insight

Building a Production-Ready Auto-Scaler for Docker Swarm on Hetzner Cloud VPS

May 29, 2026

Introduction

In the era of cloud computing, resource flexibility is paramount for maintaining application availability while optimizing infrastructure costs. While managed Kubernetes services dominate the orchestration landscape, Docker Swarm remains a compelling, lightweight choice for many enterprises due to its simplicity and low operational overhead.

However, one significant limitation of Docker Swarm compared to Kubernetes is the lack of a native, out-of-the-box cluster auto-scaler. When traffic spikes on a Virtual Private Server (VPS) cluster, engineers are often forced to manually provision new nodes—a reactive approach that risks downtime. In this comprehensive guide, we will design and implement an automated infrastructure: a custom Auto-Scaler for Docker Swarm running on Hetzner Cloud VPS, leveraging Hetzner's robust API and Prometheus metrics.

The Architecture of a Docker Swarm Auto-Scaler

Before diving into execution, it is crucial to understand how a decentralized auto-scaler functions. Unlike managed cloud environments (like AWS or GCP) where scaling is handled abstractly, a self-hosted solution on Hetzner requires three distinct layers working in harmony:

  • The Monitoring Layer: Collects real-time CPU, memory, and network utilization across all Swarm nodes using Prometheus and Node Exporter.
  • The Decision Engine: A lightweight script or service that evaluates cluster metrics against predefined thresholds (e.g., average CPU utilization exceeding 80% for more than 5 minutes).
  • The Orchestration Layer: Interacts with the Hetzner Cloud API to spin up or destroy VPS instances, and executes Docker Swarm commands to join or drain nodes seamlessly.

By decoupling these layers, we ensure that our infrastructure remains highly resilient and that scaling actions are precise, predictable, and fully automated.

Step 1: Setting Up the Infrastructure Monitoring

An auto-scaler is only as good as the metrics driving it. To prevent false positives (such as scaling up during a temporary 1-second CPU spike), we must establish a reliable metrics collection pipeline.

Deploying Prometheus and Node Exporter

We deploy Node Exporter as a global service across all existing Hetzner VPS instances. This ensures every new node automatically starts reporting its resource usage upon joining the Swarm. The configuration involves a standard Prometheus setup scraping data at 15-second intervals.

Key Metric to Watch: The primary indicator for scaling up should be a combination of cluster-wide CPU reservation and actual memory pressure, rather than a single node's performance.

Step 2: Leveraging the Hetzner Cloud API for Node Provisioning

Hetzner Cloud provides a powerful, developer-friendly REST API that allows for rapid provisioning of VPS instances (CX or CPX series). To automate this, we generate an API token from the Hetzner Cloud Console with read/write permissions.

When a scale-up event is triggered, our decision engine executes a script that performs the following sequential actions via the Hetzner API:

  1. Create Server: Request a new VPS instance using a pre-configured Ubuntu image and the appropriate SSH keys.
  2. Apply Cloud-Init: Pass a Cloud-Init script during creation to automatically install Docker, configure firewall rules, and prepare the environment.
  3. Fetch Internal IP: Retrieve the private IP address assigned to the new instance within the Hetzner Cloud Network for secure communication.

Step 3: Integrating New VPS Nodes into the Docker Swarm

Once the Hetzner VPS is active, it must securely join the existing Docker Swarm cluster without manual intervention. This is achieved by utilizing SSH remote execution or a secure configuration management tool via Cloud-Init.

The workflow follows a strict security protocol:

  • The manager node generates a temporary Swarm worker join token.
  • The auto-scaler engine securely passes this token to the newly created Hetzner instance.
  • The new instance executes the docker swarm join command over the private Hetzner network interface.
  • The manager node validates the new worker, changing its status to Ready and automatically redistributing replicated services to balance the load.

Step 4: Implementing the Scale-Down Protocol (Graceful Draining)

Scaling up protects your application's availability; scaling down protects your budget. However, abruptly terminating a VPS instance can lead to dropped requests and data corruption. A graceful scale-down strategy is non-negotiable.

When metrics indicate that resource utilization has dropped below 30% for a sustained period, the auto-scaler initiates the following teardown sequence:

  1. Drain the Node: Execute docker node update --availability drain on the Swarm manager. This instructs Docker to safely migrate all running containers off the target node onto remaining active nodes.
  2. Wait for Stabilization: Implement a cooldown period (typically 3 to 5 minutes) to ensure all connections are cleanly closed and tasks are rescheduled successfully.
  3. Leave the Swarm: Command the target instance to leave the cluster using docker swarm leave.
  4. Terminate the VPS: Issue a deletion request to the Hetzner Cloud API to stop billing for that specific instance.

Best Practices for Production Environments

Operating a custom auto-scaler in a production environment requires strict guardrails to prevent cascading failures. Implement these essential safety measures:

  • Define Hard Limits: Always set a maximum node count (e.g., max 10 nodes) in your auto-scaler configuration to prevent runaway scaling loops from inflating your Hetzner bill in the event of a DDoS attack.
  • Cool-Down Periods: Enforce a minimum time window (e.g., 10 minutes) between scaling actions to allow the cluster to stabilize before assessing metrics again.
  • State Persistence: Ensure stateful services (like databases) are pinned to dedicated, non-scaling manager nodes using placement constraints, leaving worker nodes strictly for stateless applications.

Conclusion

By building a custom auto-scaler tailored for Docker Swarm on Hetzner Cloud, you achieve the best of both worlds: the operational simplicity of Docker Swarm and the cost-effective flexibility of an automated elastic infrastructure. While it requires initial development effort, this architecture drastically reduces manual system administration and ensures your applications remain responsive under heavy loads without overpaying for idle computing power.

Building a Production-Ready Auto-Scaler for Docker Swarm on Hetzner Cloud VPS | DPTCloud