Building a Free VPS Monitoring System with Prometheus, Grafana, and Telegram Alerts
Introduction: The Need for Proactive VPS Monitoring
In today's digital landscape, virtual private servers (VPS) power countless applications, from personal projects to business-critical systems. However, without proper monitoring, you're essentially flying blind. Server downtime, performance degradation, or security incidents can occur without warning, leading to service disruption, data loss, and frustrated users. While commercial monitoring solutions exist, they often come with recurring costs and complexity that may not suit every budget or technical requirement.
This is where open-source tools shine. By combining Prometheus for metrics collection, Grafana for visualization, and Telegram for alerting, you can build a comprehensive monitoring system that's both powerful and completely free. This approach gives you full control over your monitoring infrastructure while avoiding vendor lock-in and subscription fees.
Understanding the Monitoring Stack Architecture
Before diving into implementation, it's crucial to understand how these components work together. The architecture follows a clear data flow pattern:
- Prometheus acts as the metrics collector and time-series database. It periodically scrapes metrics from your VPS and stores them efficiently.
- Node Exporter runs on each monitored VPS, exposing system metrics in a format Prometheus can understand.
- Grafana connects to Prometheus as a data source, providing rich visualization dashboards for analyzing metrics.
- Alertmanager (part of Prometheus) handles alert routing and deduplication.
- Telegram Bot receives alerts from Alertmanager and delivers them to your preferred chat or channel.
This modular architecture allows you to scale your monitoring as your infrastructure grows. You can start with a single VPS and expand to dozens without changing the fundamental setup.
Step-by-Step Implementation Guide
1. Preparing Your Monitoring Server
Begin by setting up a dedicated VPS for your monitoring stack. While you can install these components on an existing server, a separate instance ensures monitoring continues even if your application servers experience issues. A basic VPS with 1-2 GB RAM and 20 GB storage is sufficient for most small to medium deployments.
Install Docker and Docker Compose, which will simplify deployment and management:
# Update package lists
sudo apt update
# Install Docker
sudo apt install docker.io docker-compose -y
# Start and enable Docker service
sudo systemctl start docker
sudo systemctl enable docker2. Deploying Prometheus with Docker Compose
Create a docker-compose.yml file to define your monitoring services. This approach ensures consistent deployment and easy updates:
version: '3.8'
services:
prometheus:
image: prom/prometheus:latest
container_name: prometheus
restart: unless-stopped
volumes:
- ./prometheus.yml:/etc/prometheus/prometheus.yml
- prometheus_data:/prometheus
command:
- '--config.file=/etc/prometheus/prometheus.yml'
- '--storage.tsdb.path=/prometheus'
- '--web.console.libraries=/etc/prometheus/console_libraries'
- '--web.console.templates=/etc/prometheus/consoles'
- '--storage.tsdb.retention.time=30d'
ports:
- "9090:9090"
networks:
- monitoring
alertmanager:
image: prom/alertmanager:latest
container_name: alertmanager
restart: unless-stopped
volumes:
- ./alertmanager.yml:/etc/alertmanager/alertmanager.yml
- alertmanager_data:/alertmanager
ports:
- "9093:9093"
networks:
- monitoring
grafana:
image: grafana/grafana:latest
container_name: grafana
restart: unless-stopped
volumes:
- grafana_data:/var/lib/grafana
environment:
- GF_SECURITY_ADMIN_PASSWORD=YourSecurePassword123
ports:
- "3000:3000"
networks:
- monitoring
volumes:
prometheus_data:
alertmanager_data:
grafana_data:
networks:
monitoring:
driver: bridge3. Configuring Prometheus to Scrape VPS Metrics
Create a prometheus.yml configuration file that defines which servers to monitor. For each VPS, you'll need to install Node Exporter and configure Prometheus to scrape it:
global:
scrape_interval: 15s
evaluation_interval: 15s
alerting:
alertmanagers:
- static_configs:
- targets:
- alertmanager:9093
rule_files:
- "alerts.yml"
scrape_configs:
- job_name: 'prometheus'
static_configs:
- targets: ['localhost:9090']
- job_name: 'vps-nodes'
static_configs:
- targets: ['vps1-ip:9100', 'vps2-ip:9100', 'vps3-ip:9100']
metrics_path: /metrics
scheme: http4. Installing Node Exporter on Target VPS
On each VPS you want to monitor, install and run Node Exporter:
# Download the latest Node Exporter
wget https://github.com/prometheus/node_exporter/releases/download/v1.6.0/node_exporter-1.6.0.linux-amd64.tar.gz
# Extract the archive
tar xvfz node_exporter-1.6.0.linux-amd64.tar.gz
# Move binary to system directory
sudo mv node_exporter-1.6.0.linux-amd64/node_exporter /usr/local/bin/
# Create systemd service file
sudo nano /etc/systemd/system/node_exporter.serviceAdd the following service configuration:
[Unit]
Description=Node Exporter
After=network.target
[Service]
User=node_exporter
Group=node_exporter
Type=simple
ExecStart=/usr/local/bin/node_exporter
[Install]
WantedBy=multi-user.targetStart and enable the service:
sudo systemctl daemon-reload
sudo systemctl start node_exporter
sudo systemctl enable node_exporter5. Configuring Telegram Alerts
Create a Telegram bot through BotFather and note your bot token. Then configure Alertmanager to send notifications:
global:
resolve_timeout: 5m
route:
group_by: ['alertname']
group_wait: 10s
group_interval: 10s
repeat_interval: 1h
receiver: 'telegram'
receivers:
- name: 'telegram'
webhook_configs:
- url: 'https://api.telegram.org/botYOUR_BOT_TOKEN/sendMessage'
send_resolved: true
max_alerts: 5
http_config:
basic_auth:
username: 'YOUR_CHAT_ID'
password: ''6. Setting Up Alert Rules
Define meaningful alert conditions in alerts.yml:
groups:
- name: vps_alerts
rules:
- alert: HighCPUUsage
expr: 100 - (avg by(instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) > 80
for: 5m
labels:
severity: warning
annotations:
summary: "High CPU usage on {{ $labels.instance }}"
description: "CPU usage is above 80% for 5 minutes"
- alert: LowMemory
expr: (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100 < 10
for: 5m
labels:
severity: critical
annotations:
summary: "Low memory on {{ $labels.instance }}"
description: "Available memory is below 10%"
- alert: DiskSpaceCritical
expr: (node_filesystem_avail_bytes{mountpoint="/"} / node_filesystem_size_bytes{mountpoint="/"}) * 100 < 5
for: 2m
labels:
severity: critical
annotations:
summary: "Critical disk space on {{ $labels.instance }}"
description: "Disk space is below 5% on root partition"Creating Effective Grafana Dashboards
After accessing Grafana at http://your-server-ip:3000 (default credentials: admin/YourSecurePassword123), add Prometheus as a data source. Then import or create dashboards that provide actionable insights:
- System Overview Dashboard: Display CPU, memory, disk, and network usage across all servers
- Service-Specific Dashboard: Monitor application-specific metrics if you're running web servers, databases, or other services
- Alert Status Dashboard: Visualize current alert states and history
Grafana's extensive library of community dashboards provides excellent starting points. Search for "Node Exporter Full" dashboard ID 1860 for a comprehensive system monitoring template.
Security Considerations and Best Practices
While building your monitoring system, security should remain a priority:
- Firewall Configuration: Restrict access to monitoring ports (9090, 3000, 9093) to trusted IP addresses only
- Reverse Proxy: Use Nginx or Apache as a reverse proxy with SSL termination for secure external access
- Authentication: Implement strong passwords for Grafana and consider OAuth integration
- Regular Updates: Keep all components updated to patch security vulnerabilities
- Backup Strategy: Regularly backup Prometheus data and Grafana dashboard configurations
Remember: Your monitoring system is only as reliable as its own infrastructure. Consider deploying a second monitoring instance for high-availability scenarios, or at minimum, ensure regular backups of your configuration and data.
Advanced Monitoring Scenarios
Once your basic monitoring is operational, consider these enhancements:
Application-Level Monitoring
Extend beyond system metrics by monitoring your actual applications. Most modern frameworks and databases offer Prometheus exporters or native metrics endpoints:
- MySQL/MariaDB: Use mysqld_exporter
- PostgreSQL: Use postgres_exporter
- Nginx/Apache: Configure status modules and scrape with Prometheus
- Custom applications: Implement Prometheus client libraries in your code
Blackbox Monitoring
Add blackbox_exporter to monitor external service availability from your VPS perspective. This allows you to check HTTP/HTTPS endpoints, TCP ports, ICMP ping, and DNS resolution.
Log Integration
Combine Loki with Prometheus for a full observability stack. Loki provides log aggregation that correlates with your metrics, helping you troubleshoot issues more effectively.
Cost Analysis: Free vs. Commercial Solutions
Let's compare the cost of this open-source solution with popular commercial alternatives:
- Self-hosted Open Source: $0 for software + VPS hosting costs ($5-20/month)
- Datadog: Starts at $15/host/month
- New Relic: Starts at $0.30/GB ingested + $0.10/host/hour
- SolarWinds: $1,600+ for perpetual license + maintenance
For small to medium deployments, the self-hosted approach offers significant savings while providing comparable functionality. The trade-off is the time required for setup and maintenance, which decreases as you become familiar with the tools.
Maintenance and Troubleshooting Tips
Regular maintenance ensures your monitoring system remains reliable:
- Weekly: Check alert functionality by triggering test alerts
- Monthly: Review and update alert thresholds based on historical data
- Quarterly: Update all components to latest stable versions
- Biannually: Review dashboard effectiveness and retire unused panels
Common issues and solutions:
- Prometheus scraping failures: Check firewall rules and Node Exporter service status
- High memory usage: Adjust retention periods or increase VPS resources
- Missing metrics: Verify Node Exporter version compatibility
- Alert fatigue: Fine-tune alert thresholds and grouping rules
Conclusion: Taking Control of Your Infrastructure
Building a free VPS monitoring system with Prometheus, Grafana, and Telegram alerts empowers you with complete visibility into your infrastructure without ongoing software costs. This solution provides the foundation for proactive system management, allowing you to identify and address issues before they impact users.
The flexibility of open-source tools means you can extend this system as your needs evolve—adding more servers, integrating application metrics, or implementing advanced alerting logic. While the initial setup requires technical investment, the long-term benefits of reduced downtime, better performance insights, and cost savings make this approach worthwhile for any serious VPS operator.
Start with a single server, implement the basic monitoring stack, and gradually expand as you become comfortable with the tools. Within a few hours, you'll have transformed from reactive troubleshooting to proactive infrastructure management—all without spending a dime on monitoring software.
