Scaling Centralized Logging: Building a High-Performance Vector and ClickHouse Architecture for 50+ Microservices VPS
Introduction: The Logging Dilemma in Distributed Microservices
In a modern enterprise infrastructure, moving from a monolithic architecture to distributed microservices offers unparalleled scalability and agility. However, this transition introduces a significant operational challenge: log fragmentation. When your application spans more than 50 Virtual Private Servers (VPS), manually SSH-ing into individual machines to investigate errors using grep or tail becomes a bottleneck that actively harms your Mean Time to Resolution (MTTR).
For years, the standard remedy has been the ELK (Elasticsearch, Logstash, Kibana) or LGTM (Loki, Grafana, Tempo, Mimir) stacks. While powerful, these solutions often demand substantial CPU and memory footprints. Running heavy Java-based or resource-intensive logging agents on 50+ production nodes can eat into the compute resources meant for your core applications, leading to bloated cloud bills. This blog post explores a highly efficient, modern alternative for centralized logging: combining Vector as the lightweight data aggregator and ClickHouse as the ultra-fast, column-oriented analytical database.
Why Vector and ClickHouse? A Performance-Driven Evaluation
Before diving into the architecture, it is essential to understand why this specific combination outperforms traditional logging infrastructure at scale.
Vector: The Lightweight Observability Router
Developed in Rust, Vector is a high-performance observability data router. It replaces resource-heavy shippers like Logstash or Fluentd. Its core advantages include:
- Minimal Resource Footprint: Vector consumes a fraction of the memory and CPU compared to its competitors, making it safe to deploy across 50+ production VPS instances without affecting tenant applications.
- Memory Efficiency and Safety: Built on Rust, it guarantees memory safety and eliminates garbage collection pauses.
- Powerful VRL (Vector Remap Language): Allows for rich parsing, structuring, and filtering of logs at the edge before sending them to the central server.
ClickHouse: The Columnar Data Warehouse
ClickHouse is an open-source, column-oriented OLAP (Online Analytical Processing) database management system. While Elasticsearch stores data in inverted indices, ClickHouse stores data column-by-column. This design choice yields massive benefits for log management:
- Exceptional Compression Rates: Logs often contain repetitive data (IPs, status codes, service names). ClickHouse compresses this data heavily, often achieving a 5x to 10x reduction in disk usage compared to Elasticsearch.
- Blazing-Fast Analytical Queries: Searching through billions of log rows using SQL queries takes milliseconds, allowing teams to build instant dashboards and debugging reports.
- Cost Efficiency: Lower compute and storage requirements translate directly to reduced cloud infrastructure costs.
The Architecture: Centralized Logging for 50+ VPS Nodes
To implement centralized logging seamlessly across a large VPS footprint, a two-tier architecture is highly effective. The topology is divided into the Edge Layer and the Storage & Visualization Layer.
1. The Edge Layer (Production VPS Node)
Each of the 50+ VPS nodes runs its respective microservices (deployed via Docker, Kubernetes, or bare-metal binaries). Alongside these services, a lightweight Vector Agent is installed. The agent is configured to:
- Harvest logs from log files, systemd-journald, or Docker container stdout streams.
- Parse and sanitize raw strings into structured JSON formats using VRL.
- Batch and stream the structured logs securely over TLS to the centralized database cluster.
2. The Centralized Storage and Analytics Server
At the center of your infrastructure sits a dedicated, high-performance cluster or instance running ClickHouse. ClickHouse receives the structured log batches from all 50+ Vector agents simultaneously. Because ClickHouse handles concurrent inserts exceptionally well through batching, it smoothly ingests thousands of log lines per second without degrading search performance.
Key Architectural Insight: Always ensure that Vector agents perform batching before sending data to ClickHouse. ClickHouse prefers fewer, larger writes (e.g., thousands of rows at a time) rather than frequent, single-row inserts to prevent generating too many data parts on disk.
Step-by-Step Implementation Guide
Step 1: Setting up the ClickHouse Schema
First, we need to connect to our central ClickHouse instance and create a optimized table structure using the MergeTree engine family. Below is a production-ready SQL schema designed specifically for microservice logs:
CREATE TABLE system.microservice_logs (
timestamp DateTime64(3, 'UTC'),
service_name LowCardinality(String),
environment LowCardinality(String),
host String,
level LowCardinality(String),
message String,
trace_id String,
span_id String,
attributes Map(String, String)
) ENGINE = MergeTree()
PARTITION BY toYYYYMM(timestamp)
ORDER BY (environment, service_name, level, timestamp);Note: Using the LowCardinality type for fields like service_name and level optimizes storage and drastically accelerates queries because these values repeat frequently across millions of rows.
Step 2: Configuring Vector Agents on the Edge VPS
On each of your 50+ VPS nodes, install the Vector binary and create a configuration file (vector.yaml) to collect, parse, and ship logs. Here is an optimized sample configuration:
sources:
docker_logs:
type: "docker_logs"
transforms:
parse_and_structure:
type: "remap"
inputs:
- "docker_logs"
source: |
# Standardize fields using Vector Remap Language (VRL)
.timestamp = .timestamp || now()
.environment = "production"
.service_name = .label.com_docker_compose_service || .container_name
.level = upcase(string(.level)) || "INFO"
.trace_id = string(.trace_id) || ""
.span_id = string(.span_id) || ""
sinks:
clickhouse_output:
type: "clickhouse"
inputs:
- "parse_and_structure"
endpoint: "http://your-central-clickhouse-ip:8123"
database: "system"
table: "microservice_logs"
skip_unknown_fields: true
compression: "gzip"
batch:
max_bytes: 5242880 # 5MB
timeout_secs: 2Visualization and Monitoring
Once logs are securely streaming into ClickHouse, engineering teams need a way to visualize anomalies, monitor system health, and debug issues. The industry standard tool for this layer is Grafana. By installing the official ClickHouse Data Source plugin for Grafana, you can instantly build high-performance dashboards.
Engineers can write standard SQL queries directly inside Grafana to plot error rates over time, track response latencies across different microservices, or filter through raw logs based on specific trace_id attributes to trace requests seamlessly across service boundaries.
Conclusion and Best Practices
Building a centralized logging system using Vector and ClickHouse provides a scalable, enterprise-grade architecture capable of effortlessly handling the output of 50+ microservices VPS instances. By shifting away from heavy indexing engines to a columnar database approach, you achieve massive storage savings and lightning-fast search queries while minimizing the compute overhead on your production nodes.
As you transition to production, remember these critical best practices:
- Implement Strict Retention Policies: Use ClickHouse TTLs (Time-To-Live) to automatically drop or archive logs older than 14 or 30 days to save costs.
- Secure the Transport Layer: Always configure TLS/HTTPS for the communication channels between your distributed Vector agents and the centralized ClickHouse instance.
- Tune Batch Sizes: Adjust the Vector batch configuration based on your peak traffic periods to find the perfect balance between real-time visibility and database write efficiency.
