Building a $0 Distributed Tracing and Error Tracking System for Microservices using GlitchTip and Grafana Tempo on a VPS
Introduction to Cost-Effective Observability
In a modern microservices architecture, monitoring system health is no longer as simple as checking a single application log. A single user request might traverse dozens of isolated services, databases, and external APIs. When a request fails or encounters latency, identifying the exact root cause becomes a daunting needle-in-a-haystack problem. This is where Distributed Tracing and centralized error tracking become non-negotiable requirements.
However, enterprise observability platforms like Datadog, New Relic, or SaaS-hosted Sentry often come with prohibitive pricing models. As your transaction volume grows, your monitoring bill can easily skyrocket, penalizing your business for scaling. Fortunately, by leveraging open-source tools and a self-hosted Virtual Private Server (VPS), you can build a robust, production-grade telemetry pipeline for exactly $0 in licensing fees. This guide will walk you through deploying GlitchTip for exception tracking and Grafana Tempo for high-volume distributed tracing on a single VPS instance.
---Why GlitchTip and Grafana Tempo?
To build an optimal zero-cost stack, we must select tools that are both feature-rich and highly resource-efficient. Many open-source platforms demand massive clusters just to run baseline services; our selected stack is purposefully designed for lean operations.
### GlitchTip: The Lightweight Sentry AlternativeGlitchTip is an open-source, self-hosted error tracking platform that is fully compatible with Sentry's official SDKs. It aggregates errors, monitors application uptime, and tracks performance metrics. Unlike modern self-hosted Sentry—which requires significant memory and CPU overhead due to its heavy Kafka and ClickHouse dependencies—GlitchTip is written in Python/Django and utilizes a lightweight PostgreSQL backend, making it ideal for cost-constrained VPS deployments.
### Grafana Tempo: High-Scale, Low-Cost TracingGrafana Tempo is an open-source, high-scale distributed tracing backend. What makes Tempo revolutionary is its object-storage-only architecture. While traditional tracing systems like Jaeger or Zipkin historically required Elasticsearch or Cassandra to index traces, Tempo only requires basic object storage (such as MinIO, AWS S3, or local disk blocks) and generates indexes on the fly. This dramatically reduces memory utilization and storage costs, allowing you to process millions of spans on a modest VPS budget.
---System Architecture Overview
Before diving into the deployment configurations, it is critical to understand how telemetry data flows through our infrastructure. Our microservices will embed standardized, open-source SDKs to push data out asynchronously, ensuring zero performance degradation on the user-facing application path.
Architecture Flow: Application Microservices → OpenTelemetry SDK → Grafana Alloy / Grafana Agent (Optional Collector) → Grafana Tempo (Traces) & GlitchTip (Errors) → Unified Visualization via Grafana Dashboards.
By keeping both backends on a single VPS, we minimize network egress costs and latency, keeping our operational overhead strictly bounded by our baseline VPS hosting fee.
---Step-by-Step Deployment Guide via Docker Compose
We will use Docker Compose to orchestrate our entire monitoring stack. This ensures reproducibility and clean isolation of our services. Ensure your VPS has Docker and Docker Compose plugin pre-installed.
1. Setting Up the Directory Structure
Log into your VPS via SSH and create the necessary directory structure to store configurations and persistent volumes:
mkdir -p monitoring/{glitchtip,tempo,grafana}
cd monitoring2. Configuring Grafana Tempo
Create a configuration file for Tempo at ./tempo/tempo.yaml. This instructs Tempo to use the local filesystem as its object storage block, eliminating the need for an external cloud storage provider.
stream_jobs_max_bytes: 1048576
server:
http_listen_port: 3200
distributor:
receivers:
otlp:
protocols:
http:
grpc:
ingester:
lifecycler:
ring:
kvstore:
store: inmemory
replication_factor: 1
storage:
trace:
backend: local
local:
path: /var/tempo/wal
wal:
path: /var/tempo/wal
local_blocks:
path: /var/tempo/blocks3. Writing the Unified Docker Compose File
Create a docker-compose.yml file in your root monitoring/ folder. This unified file provisions PostgreSQL (shared or split via databases), GlitchTip, Grafana Tempo, and Grafana for visualization.
version: '3.8'
services:
# Database for GlitchTip
postgres:
image: postgres:15-alpine
environment:
POSTGRES_DB: glitchtip
POSTGRES_USER: postgres
POSTGRES_PASSWORD: super_secure_password
volumes:
- pgdata:/var/lib/postgresql/data
restart: unless-stopped
# GlitchTip Web & Worker
glitchtip:
image: glitchtip/glitchtip:v4.0
depends_on:
- postgres
environment:
- DATABASE_URL=postgres://postgres:super_secure_password@postgres:5432/glitchtip
- SECRET_KEY=your_production_secret_key_here
- PORT=8000
- GLITCHTIP_DOMAIN=[https://glitchtip.yourdomain.com](https://glitchtip.yourdomain.com)
ports:
- "8000:8000"
restart: unless-stopped
# Grafana Tempo (Tracing Backend)
tempo:
image: grafana/tempo:latest
command: [ "-config.file=/etc/tempo.yaml" ]
volumes:
- ./tempo/tempo.yaml:/etc/tempo.yaml
- tempodata:/var/tempo
ports:
- "4317:4317" # OTLP gRPC
- "4318:4318" # OTLP HTTP
- "3200:3200" # Tempo API
restart: unless-stopped
# Grafana (Frontend UI)
grafana:
image: grafana/grafana:latest
volumes:
- grafanadata:/var/lib/grafana
ports:
- "3000:3000"
restart: unless-stopped
volumes:
pgdata:
tempodata:
grafanadata:Deploy the stack by executing docker compose up -d. Verify all containers are running successfully using docker compose ps.
Configuring Microservices for Telemetry
With our infrastructure active, we must now configure our microservices to ship telemetry data. We will utilize industry-standard OpenTelemetry (OTel) for tracing and the standard Sentry SDK for error logging.
### 1. Error Tracking with GlitchTip SDKBecause GlitchTip uses the Sentry protocol, you can use standard Sentry packages natively. Here is a Node.js/Express framework example:
const Sentry = require("@sentry/node");
Sentry.init({
dsn: "https://[email protected]/1",
tracesSampleRate: 1.0, // Adjust in production to save bandwidth
});
// Application middleware
app.use(Sentry.Handlers.requestHandler());
app.use(Sentry.Handlers.errorHandler());### 2. Distributed Tracing with OpenTelemetryTo link downstream API calls across microservices, configure your OpenTelemetry SDK to export data directly to Tempo's OTLP receiver endpoint (http://your-vps-ip:4318/v1/traces or via gRPC on 4317). Ensure you enable W3C Trace Context propagation in your HTTP client headers so that the traceparent header automatically moves from Service A to Service B.
Unifying the Observability Experience in Grafana
One of the hidden superpowers of this specific combination is the ability to tie logs, errors, and traces together into a seamless troubleshooting workflow inside Grafana.
- Log into your Grafana instance at
http://your-vps-ip:3000(Default credentials: admin/admin). - Navigate to Connections > Data Sources and click Add data source.
- Select Tempo. Set the URL to
http://tempo:3200and save.
Now, when querying traces in the Grafana Explore tab, you can visualize full timeline waterfalls of your microservice calls. To take this a step further, you can configure GlitchTip to include Trace IDs within its error metadata. When an error is triggered, you can instantly pivot from the GlitchTip dashboard straight into Grafana Tempo to see exactly what downstream database call or API latency caused that error to occur.
---Production Best Practices for $0 Operations
Running a self-hosted monitoring system on a budget requires discipline to prevent resource exhaustion. Implement these guidelines to keep your VPS stable:
- >
- Aggressive Sampling: In production environments, do not log 100% of successful traces. Implement probabilistic sampling (e.g., sampling only 5% of HTTP 200 responses, but 100% of HTTP 5xx errors) directly within your OpenTelemetry SDK configurations.
- Data Retention Policies: Set short retention windows within your Tempo block storage configuration (e.g., 3 to 7 days). Distributed tracing is primarily used for real-time debugging and immediate post-mortems; long-term analytics should be offloaded to aggregated metrics.
- Reverse Proxy Protection: Never expose your raw database or monitoring ports directly to the public internet. Always put a lightweight reverse proxy like Nginx Proxy Manager or Caddy in front of your GlitchTip and Grafana web ports, securing them with automated, free Let's Encrypt SSL certificates.
Conclusion
Building an enterprise-grade distributed tracing and exception monitoring solution does not require thousands of dollars in monthly SaaS fees. By smartly coupling GlitchTip and Grafana Tempo on a cost-effective self-hosted VPS, you regain absolute control over your telemetry data, reduce operational costs to zero dollars in software licensing, and arm your engineering team with the deep visibility required to maintain robust microservices.
