Back to articles
Technology Insight

Building a Real-Time Distributed Logging System Using Vector.dev and ClickHouse on a VPS

June 3, 2026

Introduction: The Challenge of Distributed Logging at Scale

In modern distributed architectures, managing and analyzing application logs in real time is a critical requirement for maintaining system reliability, security, and performance. However, traditional logging stacks—often built on top of the Elastic stack (ELK)—frequently prove to be resource-intensive, complex to manage, and cost-prohibitive, especially when deployed on standard Virtual Private Servers (VPS). High memory footprints and complex JVM configurations can easily overwhelm modest infrastructure.

To overcome these challenges, a new generation of data engineering tools has emerged. By pairing Vector.dev, a lightweight, ultra-fast log aggregator written in Rust, with ClickHouse, an open-source, column-oriented database management system designed for lightning-fast online analytical processing (OLAP), engineers can construct a highly efficient, real-time distributed logging pipeline. This guide provides a comprehensive framework for architecting and deploying this modern telemetry stack directly onto a VPS.

Why Vector.dev and ClickHouse?

Before diving into the technical implementation, it is essential to understand why the combination of Vector and ClickHouse represents a paradigm shift in log management.

1. Vector.dev: The High-Performance Data Router

Vector replaces traditional collectors like Logstash or Fluentd. Its core advantages include:

  • Minimal Resource Footprint: Written in Rust, Vector operates with extremely low memory and CPU utilization, making it perfect for VPS environments where resources are constrained.
  • Memory Efficiency and Safety: Out-of-memory (OOM) crashes are virtually eliminated due to Rust's memory management model.
  • End-to-End Backpressure: Vector intelligently handles downstream bottlenecks by buffering data locally, ensuring no logs are lost during traffic spikes.

2. ClickHouse: The Ultimate Analytical Engine

ClickHouse acts as the centralized log repository. Unlike row-oriented databases, ClickHouse stores data in columns, which yields massive benefits for log analysis:

  • Exceptional Compression: Columnar storage allows identical or similar log data types to be compressed aggressively, reducing storage requirements by up to 5x to 10x compared to raw text.
  • Sub-Second Queries: Querying billions of rows for specific error codes or timestamps takes milliseconds, leveraging vectorized query execution.
  • Standard SQL Support: Teams do not need to learn custom query languages like KQL or Lucene; standard SQL is fully supported.
---

System Architecture Overview

The architecture of this real-time distributed logging system is streamlined and direct, eliminating the need for intermediate message brokers like Kafka or RabbitMQ for medium-to-large workloads on a VPS:

Data Flow: Application Logs → Vector Agent (Source) → Vector Transform (Parsing/Structuring) → ClickHouse Native/HTTP Sink → ClickHouse Analytical Tables.

By routing data straight from Vector to ClickHouse, we minimize network latency, reduce points of failure, and optimize resource usage on our VPS hosting environment.

---

Step-by-Step Deployment Guide

Step 1: Preparing ClickHouse on the VPS

First, we need to install and configure ClickHouse. For ease of maintenance and isolation, deploying via Docker Compose is highly recommended.

Create a docker-compose.yml file on your VPS:

version: '3.8'
services:
  clickhouse:
    image: clickhouse/clickhouse-server:latest
    container_name: clickhouse-server
    ports:
      - "8123:8123"
      - "9000:9000"
    volumes:
      - ./ch_data:/var/lib/clickhouse
      - ./ch_logs:/var/log/clickhouse-server
    ulimits:
      nofile:
        soft: 262144
        hard: 262144
    restart: always

Once running, connect to your ClickHouse instance using a client tool or CLI to create the structured logging table. We utilize the MergeTree engine family, which is optimized for high-volume append-only data.

CREATE DATABASE IF NOT EXISTS system_logs;

CREATE TABLE system_logs.application_logs (
    timestamp DateTime64(3, 'UTC'),
    service_name LowCardinality(String),
    environment LowCardinality(String),
    level LowCardinality(String),
    message String,
    http_status UInt16,
    duration_ms Float32,
    metadata String
)
ENGINE = MergeTree()
PARTITION BY toYYYYMM(timestamp)
ORDER BY (environment, service_name, level, timestamp);

Note: Using the LowCardinality(String) data type for fields with repeatable values (like environments or log levels) drastically optimizes storage and query performance.

Step 2: Installing and Configuring Vector.dev

Next, install Vector on the client servers or the same VPS. Vector is configured via a single vector.yaml file, divided into three core components: sources, transforms, and sinks.

sources:
  app_file_logs:
    type: "file"
    include:
      - "/var/log/apps/*.log"
    read_from: "beginning"
	ransforms:
  parse_json_logs:
    type: "remap"
    inputs:
      - "app_file_logs"
    source: |
      # Parse the raw log message as JSON
      parsed, err = parse_json(.message)
      if err != null {
        .level = "ERROR"
        .message = .message
      } else {
        # Merge parsed fields into the root context
        this = merge(this, parsed)
      }
      
      # Ensure correct types and defaults
      .timestamp = parse_timestamp(.timestamp, format: "%Y-%m-%dT%H:%M:%S%.fZ") ?? now()
      .http_status = to_int(.http_status) ?? 0
      .duration_ms = to_float(.duration_ms) ?? 0.0

sinks:
  clickhouse_output:
    type: "clickhouse"
    inputs:
      - "parse_json_logs"
    endpoint: "http://YOUR_VPS_IP:8123"
    database: "system_logs"
    table: "application_logs"
    skip_unknown_fields: true
    compression: "gzip"
    batch:
      max_events: 5000
      timeout_secs: 1

In this configuration, Vector automatically batches data up to 5,000 events or every 1 second before executing a bulk HTTP POST insert into ClickHouse, ensuring optimal write performance.

---

Performance Optimization and Best Practices

To run this pipeline efficiently on limited VPS hardware, implement the following best practices:

  • Leverage Batching: ClickHouse prefers fewer, larger inserts over many small inserts. Always tune Vector's batch.max_events and batch.timeout_secs parameters based on your ingestion volume.
  • Index Wisely: The fields specified in your ORDER BY clause dictate how data is physically sorted on disk. Place the most frequently filtered columns (e.g., environment or service_name) first.
  • Manage Data Retention: Implement Time-To-Live (TTL) policies in ClickHouse to automatically drop or compress older logs:
    ALTER TABLE system_logs.application_logs MODIFY TTL timestamp + INTERVAL 30 DAY;

Conclusion

Building a real-time distributed logging system with Vector.dev and ClickHouse offers a modern alternative to bloated traditional logging frameworks. By combining the safety and speed of Rust with the unmatched analytical power of a columnar database, business enterprises can achieve enterprise-grade telemetry visualization and analysis directly on budget-friendly VPS configurations. The resulting architecture is fast, highly maintainable, and remarkably cost-effective.

Building a Real-Time Distributed Logging System Using Vector.dev and ClickHouse on a VPS | DPTCloud