Back to articles
Technology Insight

Building a Self-Healing Application Workflow: Integrating Prometheus, n8n, and Local LLMs on a VPS

June 3, 2026

Introduction: The Paradigm of Self-Healing Infrastructure

In the contemporary digital landscape, system downtime translates directly to financial loss and diminished user trust. Traditional incident management follows a reactive paradigm: an error occurs, an alert triggers, a DevOps engineer is awoken, and manual remediation begins. While robust monitoring reduces detection time, the human element remains a bottleneck.

Enter the concept of the Self-Healing Workflow. By combining real-time telemetry, low-code automation, and the contextual reasoning of artificial intelligence, organizations can build systems that not only detect anomalies but actively diagnose and fix them. This technical guide explores how to construct a fully localized, cost-effective self-healing architecture on a Virtual Private Server (VPS) utilizing three core technologies: Prometheus for monitoring, n8n for workflow orchestration, and a Local Large Language Model (LLM) for intelligent root-cause analysis and remediation execution.

The Architecture: How the Ecosystem Fits Together

Before diving into configuration, it is essential to understand the telemetry and control plane of our self-healing ecosystem. The entire architecture operates in a closed-loop feedback system, often referred to as an OODA loop (Observe, Orient, Decide, Act):

  • Observe (Prometheus): Continuously scrapes metrics from your target applications (e.g., HTTP error rates, memory leaks, thread pool exhaustion). When a predefined threshold is crossed, Prometheus fires an alert via its Alertmanager.
  • Orient & Decide (n8n): Acts as the central nervous system. It catches the Alertmanager webhook, gathers supplementary logs, and coordinates data flow.
  • Analyze (Local LLM via Ollama): Instead of relying on rigid, hard-coded scripts for every possible error variant, n8n passes the raw error logs and system state to a localized LLM. The AI analyzes the stack trace, diagnoses the root cause, and generates the precise remediation command or script.
  • Act (n8n SSH execution): n8n receives the structured response from the LLM, validates the proposed action against a safety whitelist, executes the recovery commands on the VPS via SSH, and verifies system restoration.

Step 1: Setting Up Prometheus and Alertmanager

The foundation of self-healing is accurate, low-latency detection. Prometheus achieves this by pulling time-series metrics from your application endpoints. To implement this, we must configure both Prometheus and its companion, Alertmanager, which handles alert deduplication and webhook routing.

Prometheus Configuration

First, define the alerting rules in a file named alert.rules.yml. This rule triggers if an application returns a 5xx internal server error rate exceeding 5% over a two-minute window:

groups:
  - name: app_alerts
    rules:
    - alert: HighHttp5xxErrorRate
      expr: sum(rate(http_requests_total{status=~"5.."}[2m])) / sum(rate(http_requests_total[2m])) * 100 > 5
      for: 1m
      labels:
        severity: critical
      annotations:
        summary: "High HTTP 5xx error rate detected on {{ $labels.instance }}"
        description: "5xx errors are currently at {{ $value | printf "%.2f" }}%."

Alertmanager Configuration

Next, configure alertmanager.yml to route this specific critical alert directly to our n8n webhook endpoint:

route:
  receiver: 'n8n-webhook'
  group_by: ['alertname']
  group_wait: 10s
  group_interval: 30s
  repeat_interval: 1h

receivers:
- name: 'n8n-webhook'
  webhook_configs:
  - url: 'http://localhost:5678/webhook/v1/prometheus-alert'

Step 2: Deploying a Local LLM via Ollama on the VPS

Using cloud-based AI APIs (like OpenAI or Anthropic) for infrastructure management introduces data privacy concerns, external network dependencies, and unpredictable API costs. Hosting a highly optimized, open-source model locally on your VPS mitigates these risks entirely.

For a standard VPS environment, models like Llama-3 (8B) or Mistral (7B) quantized to 4-bits offer an exceptional balance between reasoning capability and memory footprint. We utilize Ollama for streamlined model management.

Installation and Model Pull

Install Ollama via the official shell script and launch the model:

curl -fsSL [https://ollama.com/install.sh](https://ollama.com/install.sh) | sh
ollama run llama3:8b-instruct-q4_K_M

Ollama exposes a local REST API endpoint at http://localhost:11434/api/generate, allowing n8n to communicate with the model via standard HTTP requests without heavy client libraries.

Step 3: Orchestrating the Self-Healing Logic in n8n

With our detection mechanism and intelligence engine in place, we build the orchestration workflow within n8n. The workflow is constructed using visual nodes that process data sequentially:

  1. Webhook Node: Listens for incoming POST requests from Prometheus Alertmanager. It extracts metadata such as the application name, specific error metrics, and timestamp.
  2. SSH Node (Log Gathering): Before consulting the LLM, n8n logs into the target application container or directory to fetch the last 100 lines of standard error logs using:
    docker logs --tail 100 my-app-container
  3. HTTP Request Node (LLM Consultation): This node compiles a structured prompt for Ollama. The prompt structure is vital for safety and predictability.

Engineering the Prompt for Infrastructure Safety

To ensure the LLM returns execution-ready, structured output rather than conversational prose, use strict system prompting:

System Prompt: You are an expert DevOps engineer site reliability automation tool. Analyze the provided Prometheus alert and system logs. Identify the root cause. Provide exactly ONE shell command to remediate the issue inside a JSON object format. Do not write markdown. Do not write explanations.

Expected Output Format:
{ "analysis": "Brief root cause analysis", "remediation_command": "docker restart my-app-container" }

Step 4: Implementing Guardrails and Safe Execution

Allowing an AI model to execute arbitrary shell commands on your production VPS carries inherent operational risks. A robust self-healing system requires strict guardrails. Never wire the LLM output directly into an SSH execution node.

The Validation Matrix

Within n8n, insert a Switch Node or a Code Node (JavaScript) to act as an execution firewall. Implement the following verification checks:

  • Command Whitelisting: Match the LLM's generated remediation_command against a strict Regular Expression pattern. Allow only deterministic operational commands such as docker restart [a-zA-Z0-9_-]+, systemctl restart [a-zA-Z0-9_-]+, or clearing a specific cache directory.
  • Destructive Command Blacklisting: Reject any payload containing substrings like rm -rf, chmod, chown, shutdown, or raw database drops.
  • Circuit Breakers: Track healing attempts in a localized key-value store or n8n internal variables. If an application requires a self-healing action more than twice within a 30-minute window, trip the circuit breaker, halt automated execution, and escalate to an engineer via Slack or PagerDuty.

Step 5: Verifying System Recovery and Notifications

Once the validated remediation command is safely executed via the n8n SSH Node, the workflow must not simply assume success. It transitions into a verification phase:

  1. Wait/Delay Node: Pause execution for 30 to 45 seconds to allow the application context to re-initialize and stabilize.
  2. HTTP Request Node (Health Check): Perform a targeted probe against the application’s /health or /ready endpoint.
  3. Conditional If Node: Check if the status code returns 200 OK.

Finally, utilize an integration node (such as Telegram, Slack, or Discord) to publish a summary report to the engineering team. This ensures full operational visibility, providing details on what broke, why the LLM believed it broke, what action was taken, and whether the system successfully recovered.

Conclusion: Embracing Autonomous SRE

By marrying the precision of Prometheus monitoring, the flexible glue of n8n, and the semantic reasoning of local LLMs, you successfully transition your infrastructure from a reactive state to an autonomous, self-healing architecture. Operating this entire stack locally on your VPS guarantees data privacy, minimizes operational overhead, and drives recovery time objectives (RTO) down from hours to mere seconds. As open-source models continue to mature, the frontier of automated site reliability engineering is accessible to teams of any scale.

Building a Self-Healing Application Workflow: Integrating Prometheus, n8n, and Local LLMs on a VPS | DPTCloud