Back to articles
Technology Insight

Scaling Web Scraping Infrastructure: Building a Distributed Scrapy Cluster and Crawllee System Across 5 Cheap ARM VPS with Redis

May 27, 2026

Introduction: The Challenge of Scale vs. Budget in Web Ingestion

In the data-driven business landscape, web scraping has evolved from simple scripts into enterprise-level data ingestion infrastructure. However, as organizations scale their data requirements, they quickly encounter a dual challenge: the computational bottleneck of single-machine crawlers and the skyrocketing infrastructure costs of traditional cloud compute instances. Scaling horizontally is the logical answer, but doing so efficiently requires a smart architectural blueprint.

This technical guide demonstrates how to build a highly scalable, resilient, and production-grade distributed web scraping system. By leveraging 5 budget-friendly ARM-based Virtual Private Servers (VPS), coordinating them via Redis as a centralized message broker, and combining the robust orchestration of Scrapy Cluster with the modern, anti-bot handling capabilities of Crawllee, you can achieve enterprise-scale scraping at a fraction of the traditional cost.

Why ARM Architecture and This Hybrid Stack?

Before diving into the configuration, it is essential to understand why this specific hardware and software combination provides a massive competitive advantage for modern data engineering pipelines.

  • Unmatched Cost-to-Performance Ratio: Modern ARM-based VPS offerings (such as those from Hetzner, Oracle Cloud Free Tier, or AWS Graviton) offer significantly higher core counts and RAM allocations per dollar compared to traditional x86_64 instances. Web scraping is inherently I/O-bound and highly concurrent; more cores mean more simultaneous workers.
  • Scrapy Cluster for Orchestration: Scrapy is the gold standard for structured data extraction. By utilizing Scrapy Cluster, we move away from isolated spiders and transition into a stateful, multi-tenant ecosystem that can distribute jobs dynamically across nodes.
  • Crawllee for Modern Web Ingestion: Many modern websites heavily rely on JavaScript rendering and implement aggressive anti-bot features. Crawllee (Node.js/Python) excels at browser finger-printing mimicry, stealth headless browser management, and automated proxy rotation, making it the perfect complementary worker for Javascript-heavy targets.

Architectural Blueprint: The Distributed Ecosystem

To orchestrate a 5-node cluster efficiently without nodes overwriting each other or duplicating effort, we decouple the components into a hub-and-spoke model centralized around a Redis cluster.

Our infrastructure allocation for the 5 ARM VPS nodes is structured as follows:

  1. Node 1 (The Control Plane & Broker): Hosts the primary Redis instance (Hiredis optimized), the Scrapy Cluster Kafka/Redis monitor, and API endpoints for job submissions.
  2. Nodes 2 & 3 (Scrapy Cluster Workers): Optimized for high-throughput, structured API and static HTML scraping using Python.
  3. Nodes 4 & 5 (Crawllee Advanced Workers): Process JavaScript-heavy sites utilizing headless Chromium/Playwright instances managed by Crawllee.
Key Architectural Principle: Workers must remain entirely stateless. All crawling states, duplicated URL filters (Bloom filters), and request queues reside strictly within the centralized Redis layer on Node 1. This ensures that if Node 4 crashes due to a memory leak in a headless browser, the cluster redistributes the pending queues immediately.

Setting Up the Centralized Redis Queue on ARM

First, we must configure Node 1 to handle high-concurrency connections from the remaining four worker nodes. Since ARM architecture benefits from optimized memory bandwidth, we compile or install the latest Redis stable version and tune its configuration file (redis.conf).

Essential configuration adjustments for the Redis coordinator include:

  • bind 0.0.0.0 - Ensuring the interface listens to internal private VPC network IPs securely.
  • maxmemory 4gb - Allocating adequate memory for deep crawling queues, utilizing a volatile-lru eviction policy.
  • protected-mode yes - Mandating strong internal password authentication.

By leveraging Scrapy-Redis integration within Scrapy Cluster, default memory queues are replaced with Redis Spider Queues (Priority, FIFO, or LIFO) and a Redis-backed duplicate filter, preventing nodes from wasting bandwidth on already processed pages.

Implementing Scrapy Cluster Workers (Nodes 2 & 3)

On the Python-centric nodes, Scrapy Cluster acts as a long-running daemon. Unlike standard Scrapy scripts that terminate upon job completion, these workers listen continuously to the Redis queue for incoming JSON commands.

A typical job payload pushed to the Redis queue looks like this:

{ "spiderid": "e_commerce_spider", "url": "[https://example.com/products](https://example.com/products)", "crawlid": "campaign_q2_2026", "appid": "dashboard_analytics" }

The worker nodes instantly pick up the job, process the target domain, extract data points according to pre-defined selectors, and stream the scraped items out to a centralized database or an S3-compatible object storage bucket, leaving the local VPS disk footprint virtually at zero.

Integrating Crawllee for Stealth and JavaScript Execution (Nodes 4 & 5)

Static scraping fails when encountering highly interactive web applications or advanced Cloudflare/Akamai protective walls. This is where Nodes 4 and 5, powered by Crawllee, enter the workflow. Crawllee reads its targets from the very same Redis architecture by utilizing custom Redis client adapters.

Crawllee provides critical advantages in this hybrid system:

  • Automatic Proxy Management: Intelligently rotates proxies and monitors their health, retiring banned IPs dynamically.
  • Human-like Fingerprints: Automatically alters HTTP headers, TLS signatures, and navigator object traits to blend in with legitimate user traffic.
  • Concurrency Scaling: Dynamically adjusts the number of open headless browser tabs based on the available ARM CPU and memory capacity of the VPS, preventing system crashes.

Monitoring, Error Handling, and Resilience

Deploying a distributed system without comprehensive monitoring is flying blind. When managing 5 low-cost instances, hardware throttling or network drops are expected events rather than rare exceptions. To maintain an optimized pipeline, implement the following DevOps strategies:

1. Distributed Logging

Do not store logs locally on the ARM instances. Use a lightweight logging shipper (like Fluent Bit) on each node to stream error logs and warnings to a centralized visualization platform.

2. Graceful Backoff and Circuit Breakers

If a specific web target begins returning 429 Too Many Requests or 503 Service Unavailable, the cluster coordinator must automatically pause the specific domain's queue within Redis. This prevents the cluster from burning through your proxy pools and flags the site for architectural adjustment.

3. Memory Recycling

Headless browsers on Nodes 4 and 5 will inevitably leak memory over extended operations. Implement a Cron utility or container restart policy (via Docker Compose) to recycle Crawllee containers once every 6 hours to flush dead processes and optimize ARM memory structures.

Conclusion: Scalable Data Infrastructure Within Reach

Building a robust, enterprise-grade web scraping pipeline no longer requires high-tier, expensive cloud instances. By pairing the processing efficiency of low-cost ARM VPS architecture with the smart orchestration of Scrapy Cluster, the advanced rendering power of Crawllee, and the lightning-fast queuing of Redis, you create a system that is both incredibly powerful and fiscally optimized.

This decentralized approach provides your business with the agility to scale data operations up or down dynamically, ensuring your data pipelines remain populated with high-quality, real-time web insights while maintaining strict infrastructure cost control.

Scaling Web Scraping Infrastructure: Building a Distributed Scrapy Cluster and Crawllee System Across 5 Cheap ARM VPS with Redis | DPTCloud