Back to articles
Technology Insight

Building a Comprehensive Monitoring System for Multiple VPS Servers using Prometheus and Grafana

April 14, 2026
Comprehensive VPS Monitoring with Prometheus and Grafana 2026

Building a Comprehensive Monitoring System for VPS Clusters using Prometheus and Grafana

When managing a large number of Virtual Private Servers (VPS), relying solely on the basic resource charts provided by a provider's Control Panel is insufficient. Control Panels often suffer from high latency and lack depth regarding specific metrics like Disk I/O, system errors, or custom application metrics. To ensure your system operates 24/7/365, you need a centralized monitoring solution capable of real-time alerting and detailed historical data storage. The duo of Prometheus and Grafana remains the gold standard in the DevOps industry in 2026 to solve this challenge.

1. Centralized Monitoring Architecture

Modern monitoring architecture is based on the "Pull" model, where Prometheus actively retrieves data from target servers. This reduces the overhead on the monitored VPS and allows administrators to easily manage the server list.

  • Prometheus Server: The heart of the system, responsible for collecting and storing data in a Time-series format.
  • Node Exporter: A lightweight "Agent" installed on each VPS to extract hardware metrics from the Linux kernel.
  • Grafana: The visualization interface that turns raw numbers into dynamic charts and professional dashboards.
  • Alertmanager: The component that handles alerts, sending notifications via Telegram, Slack, or Email when issues occur.

// Data structure simulating the configuration of Targets (VPS) in the Monitoring system
interface MonitoringTarget {
    host: string;
    port: number;
    labels: {
        env: 'production' | 'staging';
        service: string;
        region: string;
    };
}

const vpsTargets: MonitoringTarget[] = [
    { host: '123.45.67.89', port: 9100, labels: { env: 'production', service: 'web-backend', region: 'Hanoi' } },
    { host: '123.45.67.90', port: 9100, labels: { env: 'production', service: 'database-sql', region: 'Singapore' } }
];

function generatePromConfig(targets: MonitoringTarget[]) {
    console.log("Generating prometheus.yml static_configs...");
    // Logic for automated configuration file generation
    return targets.map(t => `${t.host}:${t.port}`);
}
    

2. Installing Node Exporter on Target VPS

Node Exporter is the most critical component on each monitored VPS. It acts as a bridge between the Linux operating system and Prometheus. Its strength lies in its extremely low resource footprint (consuming only about 10-20MB of RAM) while collecting hundreds of different metrics.

To install, simply download the appropriate distribution from the Prometheus website, extract it, and run it as a Systemd service. Once running, Node Exporter opens port 9100 so Prometheus can "pull" the data.


// Logic simulating a health check for the Node Exporter service on Linux
interface ServiceStatus {
    name: string;
    isActive: boolean;
    uptime: string;
    memoryUsage: number; // in MB
}

function checkNodeExporterHealth(status: ServiceStatus): string {
    if (status.isActive && status.memoryUsage < 50) {
        return `Service ${status.name} is healthy. Uptime: ${status.uptime}`;
    }
    return `Warning: ${status.name} may have issues or high resource usage!`;
}

const currentStatus: ServiceStatus = {
    name: "node_exporter",
    isActive: true,
    uptime: "15 days",
    memoryUsage: 12.5
};
console.log(checkNodeExporterHealth(currentStatus));
    

3. Configuring Prometheus Server for Data Collection

Once Node Exporter is installed on the VPS, you need to declare their IP addresses in the prometheus.yml configuration file. This is where you define the Scrape Interval.

Pro Tip: For Production systems, set the scrape_interval between 15 and 30 seconds. Setting it too short (e.g., 1 second) will bloat the Prometheus server's storage rapidly, while setting it too long may cause you to miss sudden spikes in resource usage.

Configuration Parameter Recommended Value Purpose
global.scrape_interval 15s - 30s How often Prometheus pulls data from the VPS.
storage.tsdb.retention.time 15d - 30d How long historical data is kept in the DB.
job_name "vps-nodes" Identifier for the group of servers being monitored.

4. Advanced Data Visualization with Grafana Dashboards

Prometheus stores data exceptionally well, but its default interface is very basic. This is where Grafana shines. Instead of designing charts from scratch, you can use template Dashboards (e.g., ID: 1860) to have a professional monitoring suite ready in minutes.

In Grafana, you can observe critical parameters that Control Panels often overlook:

  • Disk I/O Latency: The delay in reading/writing to the disk. If this is high, the VPS will "hang" even if CPU usage is low.
  • Network Throttling: Check if the VPS is being throttled by the provider's bandwidth limits.
  • Context Switches: Helps determine if the OS is overloaded managing processing threads.

// Alert Threshold calculation logic for Grafana
interface AlertRule {
    metric: string;
    currentValue: number;
    threshold: number;
    duration: string; // e.g., "5m"
}

function evaluateAlert(rule: AlertRule): boolean {
    const isViolated = rule.currentValue > rule.threshold;
    if (isViolated) {
        console.warn(`ALERT: ${rule.metric} exceeded ${rule.threshold} for ${rule.duration}!`);
    }
    return isViolated;
}

const ramAlert: AlertRule = { 
    metric: "node_memory_usage_bytes", 
    currentValue: 0.92, // 92%
    threshold: 0.90, // 90%
    duration: "5m" 
};
evaluateAlert(ramAlert);
    

5. Setting Up Real-time Alerts via Telegram

Monitoring without alerting is just "watching for fun." A true monitoring system must "scream" when an issue occurs via Alertmanager.

  • CPU Alert: If CPU > 90% continuously for 5 minutes.
  • Disk Alert: If free disk space is < 10% (Prevents database corruption).
  • Node Down Alert: If a VPS completely loses connection to the Prometheus server.

Using the Telegram Bot API is the fastest and free way to receive alerts on your phone anytime, anywhere.

6. Performance Optimization and Retention

Monitoring data (Metrics) can consume a lot of storage if you monitor hundreds of VPS simultaneously. Prometheus uses efficient Time-series compression, but you still need optimization strategies:

Don't monitor everything. Use relabel_configs to drop unnecessary metrics (e.g., virtual network card parameters that aren't used). Also, split monitoring jobs by geographic Region to manage configurations more easily.


// Function to estimate the storage space required for Prometheus
function estimateStorageMB(nodes: number, metricsPerNode: number, retentionDays: number): number {
    const samplesPerDay = (24 * 60 * 60) / 15; // 15s/sample
    const bytesPerSample = 2; // Average 2 bytes
    const totalSamples = nodes * metricsPerNode * samplesPerDay * retentionDays;
    return Math.round((totalSamples * bytesPerSample) / (1024 * 1024));
}

const estimatedStorage = estimateStorageMB(10, 500, 30);
console.log(`Estimated storage for 10 VPS over 30 days: ${estimatedStorage} MB`);
    

7. Conclusion: Monitoring System Checklist

To finalize your multi-VPS monitoring system, ensure you have completed the following steps:

  1. All target VPS have Firewall port 9100 opened for the Prometheus server's IP.
  2. The prometheus.yml file has been syntax-checked before reloading.
  3. Grafana has successfully connected to Prometheus as a DataSource.
  4. Alertmanager has been tested and successfully sends messages to Telegram.
  5. The Retention policy matches the monitoring server's disk capacity.

We hope this guide helps you build a powerful "command center," allowing you to master your VPS infrastructure like a professional!