Optimizing Distributed Tracing: Deploying OpenTelemetry and SigNoz on Resource-Constrained VPS for Go and Rust Microservices
Introduction: The Observability Dilemma for Lean Infrastructure
In the era of microservices, distributed tracing has transitioned from a luxury to an absolute necessity. Understanding how requests flow through a complex web of interconnected services is critical for maintaining performance and debugging production issues. However, traditional observability stacks—often built on heavy Java-based backends or expensive SaaS platforms—frequently demand substantial computational resources or budget.
For engineering teams operating on resource-constrained Virtual Private Servers (VPS), this creates a significant challenge. Fortunately, the convergence of the OpenTelemetry (OTel) standard, SigNoz (a lightweight, ClickHouse-powered observability platform), and high-performance languages like Go and Rust opens up a new paradigm. This blog post provides a comprehensive architectural blueprint for deploying a production-ready, low-footprint distributed tracing system on low-spec infrastructure.
Why Go, Rust, OpenTelemetry, and SigNoz?
When hardware resources are limited (e.g., a VPS with 2 vCPUs and 4GB RAM or less), every megabyte of memory and CPU cycle counts. The technologies chosen for this stack complement each other perfectly to minimize overhead:
- Go and Rust: Both languages compile to native binaries, feature minimal memory footprints, and offer predictable performance. Unlike runtime-heavy languages, they do not require massive heaps, making them ideal for constrained environments.
- OpenTelemetry: As a CNCF incubating project, OTel provides a unified, vendor-agnostic standard for collecting telemetry data, eliminating vendor lock-in while maintaining high efficiency through optimized SDKs.
- SigNoz: Built on top of ClickHouse (a columnar database renowned for its blazing-fast analytical queries and high data compression ratios), SigNoz consumes a fraction of the resources required by Elasticsearch-based alternatives like the ELK stack or Jaeger with an ES backend.
Architectural Overview for Low-Spec VPS
To succeed on a low-configuration VPS, a standard "out-of-the-box" deployment will not suffice. We must strategically design the telemetry pipeline to prevent resource exhaustion.
Rule of thumb for low-spec deployments: Offload as much processing as possible from the application runtime, aggressively sample traces, and constrain the database memory usage.
Instead of applications sending traces directly to the SigNoz backend, we utilize a localized architecture:
- Application Layer (Go/Rust): Emits spans via OTel SDKs using ultra-lightweight memory exporters.
- Local OTel Collector (Agent Mode): Runs as a sidecar or a single lightweight daemon on the VPS. It receives data locally, batches it, and handles retries.
- SigNoz & ClickHouse Storage: The centralized engine that ingests, indexes, and visualizes the trace data.
Step-by-Step Optimization Guide
1. Tightening the ClickHouse & SigNoz Memory Footprint
By default, ClickHouse attempts to utilize available system memory to speed up queries. On a low-spec VPS, this can trigger the Linux Out-Of-Memory (OOM) killer. You must explicitly cap its memory in the users.xml or Docker Compose configuration:
1073741824
2
Additionally, adjust SigNoz's retention policies early. Keeping data for only 3 to 5 days drastically reduces storage overhead and keeps ClickHouse index sizes manageable within limited RAM.
2. Implementing Head-Based Sampling in Go and Rust
Tracing 100% of requests is unsustainable on minimal infrastructure. Implementing head-based sampling at the application level ensures that only a controlled percentage of traces are generated, saving CPU and network I/O.
In Go, configure your tracer provider with a ratio-based sampler:
import (
"go.opentelemetry.io/otel/sdk/trace"
)
tp := trace.NewTracerProvider(
trace.WithSampler(trace.ParentBased(trace.TraceIDRatioBased(0.10))), // Sample 10%
// ... other options
)In Rust, using the opentelemetry_sdk crate, the configuration mirrors this logic:
use opentelemetry_sdk::trace::{config, Sampler, TracerProvider};
let provider = TracerProvider::builder()
.with_config(config().with_sampler(Sampler::TraceIdRatioBased(0.10)))
.build();3. Tuning the OpenTelemetry Collector
The OTel Collector is the unsung hero of low-resource observability. To prevent spikes, fine-tune the batch processor in your otel-collector-config.yaml. Batching reduces the number of outgoing network calls to SigNoz:
processors:
batch:
send_batch_size: 1000
timeout: 10s
send_batch_max_size: 1500
memory_limiter:
check_interval: 1s
limit_percentage: 75
spike_limit_percentage: 15The memory_limiter processor is mandatory here; it drops data dropped drops safely if the collector approaches its memory ceiling, ensuring your core business applications remain unaffected.
Real-world Performance Expectations
When properly tuned, what does the resource consumption look like on a 2 vCPU / 4GB RAM VPS hosting a few Go/Rust microservices?
- Go/Rust Services: CPU overhead remains negligible (<2%). Memory allocation increases by merely 15-30MB per instance due to the OTel SDK buffering.
- OTel Collector: Consistently consumes less than 50MB of RAM.
- SigNoz + ClickHouse: Steady state memory sits around 1.2GB to 1.5GB. Under peak ingestion, ClickHouse CPU usage may spike briefly but stabilizes quickly due to the tuned thread constraints.
Conclusion
Achieving comprehensive distributed tracing does not require expensive enterprise plans or heavy, unmanageable infrastructure. By pairing the mechanical sympathy of Go and Rust with the efficiency of OpenTelemetry and the analytical power of SigNoz (ClickHouse), you can run a robust monitoring ecosystem on a budget friendly VPS. By enforcing strict memory limits, leveraging batching, and utilizing intelligent sampling, you guarantee deep visibility into your systems without sacrificing application stability.
