Back to articles
Technology Insight

Building an Ultra-Lightweight Centralized Log Observability System with Vector.dev and ClickHouse on a 4GB RAM VPS

June 3, 2026

Introduction: The Observability Dilemma for Lean Infrastructure

In modern software architecture, centralized logging is no longer a luxury—it is a critical operational necessity. However, for startups, small-to-medium enterprises (SMEs), and independent developers, traditional log aggregation stacks present a financial and technical hurdle. The industry-standard Elasticsearch, Logstash, and Kibana (ELK) stack, while powerful, is notoriously resource-hungry. Deploying ELK typically requires significant memory allocations, often making it impossible to run effectively on low-cost infrastructure.

But what if you could achieve enterprise-grade log ingestion, storage, and querying capabilities on a single Virtual Private Server (VPS) with just 4GB of RAM? By swapping resource-heavy legacy tools for a modern, high-performance alternative—specifically, Vector.dev for log routing and ClickHouse for columnar data storage—you can build an ultra-lightweight, blazing-fast observability pipeline. This guide provides a comprehensive blueprint for architecting and deploying this cost-efficient log telescope system.

Why Vector and ClickHouse? Breaking Down the Architecture

To understand why this combination works perfectly on constrained hardware, we must look at the design philosophy of both open-source tools.

Vector.dev: The High-Performance Log Agent

Written entirely in Rust, Vector is a lightweight, ultra-fast tool for building data pipelines. Unlike Logstash or Fluentd, which run on the JVM or Ruby runtimes and consume hundreds of megabytes of RAM out of the box, Vector operates with a minimal memory footprint (often under 30MB) and provides strict memory safety. It excels at collecting logs from various sources, transforming them on the fly, and routing them to upstream sinks.

ClickHouse: The Column-Oriented Database Engine

ClickHouse is an open-source, column-oriented OLAP (Online Analytical Processing) database management system. Traditional databases like MySQL or PostgreSQL store data in rows, which is ideal for transactional systems but highly inefficient for log analytics. ClickHouse stores data in columns, allowing it to execute analytical queries over billions of rows in milliseconds. Furthermore, ClickHouse offers exceptional data compression ratios (often exceeding 5:1), drastically reducing disk space requirements.

Together, Vector and ClickHouse form a synergistic pair: Vector handles the lightweight parsing and transmission, while ClickHouse stores and indexes millions of log lines using only a fraction of the hardware resources required by Elasticsearch.

System Architecture Overview

The architecture of our ultra-lightweight log monitoring system follows a streamlined, linear data flow designed to minimize CPU context switching and memory allocation:

  • Log Sources: Application logs (Nginx, Docker, systemd, or custom application JSON logs) generated on the server.
  • Vector (Agent & Aggregator): Consumes raw logs, parses them into structured JSON, filters out unnecessary noise, and batches them for efficiency.
  • ClickHouse (Storage Engine): Receives the structured batches from Vector and writes them to highly optimized, compressed columnar tables.
  • Visualization Layer: Lightweight UI dashboards (such as Grafana or ClickHouse-Keen) connect directly to ClickHouse via standard SQL queries to display real-time insights.

Step-by-Step Deployment Guide on a 4GB VPS

Before beginning the installation, ensure your VPS is running a clean installation of an LTS Linux distribution, such as Ubuntu 22.04 or 24.04, and that Docker is installed to simplify service management.

Step 1: Optimizing the OS for Low-Memory Footprints

To ensure stability on a 4GB RAM system, configuring a swap file is highly recommended to handle unexpected memory spikes without triggering the Linux Out-Of-Memory (OOM) killer.

# Create a 4GB swap file
sudo fallocate -l 4G /swapfile
sudo chmod 600 /swapfile
sudo mkswap /swapfile
sudo swapon /swapfile
# Make it persistent across reboots
echo '/swapfile none swap sw 0 0' | sudo tee -a /etc/fstab

Step 2: Configuring and Deploying ClickHouse

We will deploy ClickHouse using Docker. Because we are operating on a 4GB RAM limits, we will modify ClickHouse configuration parameters to restrict its maximum memory usage, ensuring it leaves enough headroom for the operating system and Vector.Create a docker-compose.yml file and add the following service definition:

version: '3.8'
services:
  clickhouse:
    image: clickhouse/clickhouse-server:latest
    container_name: clickhouse-server
    ports:
      - "8123:8123"
      - "9000:9000"
    volumes:
      - ./clickhouse/data:/var/lib/clickhouse
      - ./clickhouse/log:/var/log/clickhouse
    ulimits:
      nofile:
        soft: 262144
        hard: 262144
    deploy:
      resources:
        limits:
          memory: 2G

Next, we must create the database schema. Access the ClickHouse client and execute the following SQL statement to create a optimized structured table for logs:

CREATE DATABASE IF NOT EXISTS logging;

CREATE TABLE IF NOT EXISTS logging.application_logs (
    timestamp DateTime64(3, 'UTC'),
    service_name LowCardinality(String),
    environment LowCardinality(String),
    level LowCardinality(String),
    message String,
    attributes Map(String, String)
)
ENGINE = MergeTree()
PARTITION BY toYYYYMM(timestamp)
ORDER BY (environment, service_name, level, timestamp);

Note: Using the LowCardinality(String) data type for fields like level or service_name optimizes storage and query speeds by assigning internal numeric IDs to repetitive strings.

Step 3: Configuring Vector for Log Collection and Ingestion

With ClickHouse ready, we configure Vector via its native configuration file, vector.yaml. Vector will monitor local log sources, parse them, and stream them directly into our ClickHouse instance using the native HTTP interface.

sources:
  docker_logs:
    type: "docker_logs"

transforms:
  parse_logs:
    type: "remap"
    inputs:
      - "docker_logs"
    source: |
      .service_name = .container_name
      .environment = "production"
      .level = .status || "info"
      .attributes.image = .container_image
      .attributes.stream = .stream
      del(.container_name)
      del(.container_image)
      del(.status)
      del(.stream)

sinks:
  clickhouse_out:
    type: "clickhouse"
    inputs:
      - "parse_logs"
    endpoint: "http://clickhouse-server:8123"
    database: "logging"
    table: "application_logs"
    skip_unknown_fields: true
    compression: "gzip"
    batch:
      max_bytes: 5242880
      timeout_secs: 5

This configuration enforces a batching mechanism. Instead of sending a separate HTTP request for every single log line, Vector holds logs in a memory buffer until either 5MB of data is reached or 5 seconds have elapsed. This dramatically reduces the I/O load on our ClickHouse server.

Performance Evaluation & Resource Usage

Once operational, the resource efficiency of this architecture becomes clearly visible compared to traditional ELK setups:

Component Memory Usage (ELK Stack) Memory Usage (Vector + ClickHouse)
Shipper / Collector ~300MB - 500MB (Logstash) ~25MB - 50MB (Vector)
Storage & Engine ~2GB - 4GB+ (Elasticsearch) ~1GB - 1.5GB (ClickHouse)
Total Minimum RAM > 4GB (Unstable on 4GB VPS) ~1.5GB - 2GB (Highly Stable)

On a standard 4GB RAM VPS, this pipeline leaves approximately 2GB of memory completely free for hosting actual business applications, websites, or API gateways on the exact same node.

Best Practices for Maintaining the System on Restricted Hardware

To guarantee long-term stability and prevent the system from exhausting hardware resources, keep the following operational strategies in mind:

  1. Implement Strict TTLs (Time-To-Live): Log data naturally loses value over time. Configure ClickHouse to automatically purge old logs using the TTL feature:
    ALTER TABLE logging.application_logs MODIFY TTL timestamp + INTERVAL 14 DAY;
  2. Monitor Disk I/O: Because ClickHouse compresses data aggressively, it trades minor CPU overhead for disk space savings. Ensure your VPS uses SSD storage to maintain high-speed analytical queries.
  3. Tune Vector Buffering: If your application experiences massive traffic spikes, configure Vector to use a disk-backed buffer rather than a memory buffer to safeguard against out-of-memory errors during high-load scenarios.

Conclusion

Building a robust observability stack does not require a massive budget or complex cloud architectures. By stepping away from heavy JVM-based software and embracing high-efficiency modern alternatives like Vector.dev and ClickHouse, you can deploy a scalable, powerful, centralized log management platform on a budget-friendly 4GB RAM VPS. This setup provides deep operational visibility into your applications without sacrificing your infrastructure's computing power or your bottom line.

Building an Ultra-Lightweight Centralized Log Observability System with Vector.dev and ClickHouse on a 4GB RAM VPS | DPTCloud