Back to articles
Technology Insight

Self-Healing AI Infrastructure: Automated Incident Remediation with OpenTelemetry, Robusta & Local LLMs

August 21, 2026

Self-Healing AI Infrastructure: Automated Incident Remediation with OpenTelemetry Spans, Robusta, and K8s Local LLM Agents

Modern enterprise Kubernetes clusters generate immense telemetry volume. However, traditional Mean Time To Detection (MTTD) and Mean Time To Resolution (MTTR) are severely bottlenecked by human cognitive bandwidth. When an incident occurs—such as a cascading connection pool exhaustion or a memory leak in a microservice—SREs must manually correlate OpenTelemetry traces, pull pod logs, analyze Kubernetes events, and execute patch maneuvers.

This guide presents an enterprise-grade architectural blueprint for Self-Healing Kubernetes Infrastructure. By combining OpenTelemetry (OTel) trace context, Robusta.dev incident automation, and an in-cluster Local Large Language Model (LLM) agent (e.g., DeepSeek-R1 or Llama-3 running on vLLM), we build a closed-loop remediation pipeline that autonomously analyzes root causes and executes deterministic, guarded repairs without sending sensitive telemetry outside your security perimeter.

1. Architectural Overview & Workflow

The self-healing architecture decouples trace telemetry generation, alert enrichment, intelligent reasoning, and safe execution guardrails.

+-----------------------------------------------------------------------------------+
|                                KUBERNETES CLUSTER                                 |
|                                                                                   |
|  +--------------------+      OTLP      +---------------------------------------+  |
|  | Application Pods   | -------------> | OpenTelemetry Collector               |  |
|  | (OTel Instrumented)|                | (Tail-sampling & Span Metrics)        |  |
|  +--------------------+                +---------------------------------------+  |
|            |                                              |                       |
|        Failing                                      Triggers Alert                |
|        Request                                            |                       |
|            v                                              v                       |
|  +--------------------+                        +-------------------------------+  |
|  | K8s Events / Logs  | <--------------------- | Prometheus Alertmanager       |  |
|  +--------------------+                        +-------------------------------+  |
|                                                           |                       |
|                                                     Alert Payload                 |
|                                                           v                       |
|                                                +-------------------------------+  |
|                                                | Robusta Operator              |  |
|                                                | (Enrichment Engine)           |  |
|                                                +-------------------------------+  |
|                                                           |                       |
|                                                 Span Context + Logs               |
|                                                           v                       |
|                                                +-------------------------------+  |
|                                                | In-Cluster LLM Remediation    |  |
|                                                | Agent (vLLM / Ollama + API)   |  |
|                                                +-------------------------------+  |
|                                                           |                       |
|                                                  Strict JSON Action               |
|                                                           v                       |
|                                                +-------------------------------+  |
|                                                | Robusta Remediation Runner    |  |
|                                                | (RBAC & Guardrail Enforcer)   |  |
|                                                +-------------------------------+  |
|                                                           |                       |
|                                                   Executes API Patch              |
|                                                           v                       |
|                                                +-------------------------------+  |
|                                                | Kubernetes API Server         |  |
|                                                +-------------------------------+  |
+-----------------------------------------------------------------------------------+

Remediation Flow Lifecycle

  1. Telemetry Capture & Context Extraction: Application pods emit spans via OTLP. The OTel Collector identifies error spans (e.g., HTTP 5xx, DB timeout) and attaches contextual IDs (`trace_id`, `span_id`).
  2. Enrichment Trigger: Prometheus or OTel Collector metrics fire an alert to Prometheus Alertmanager, which forwards it to the Robusta engine.
  3. Deep Enrichment: Robusta queries the Jaeger/Tempo API for full trace spans leading up to the crash and collects the last 100 lines of pod stdout/stderr alongside Kubernetes events.
  4. Local LLM Analysis: Robusta sends the structured incident context to a local, isolated LLM inference service (vLLM running a quantized model). The model outputs a strictly structured JSON remediation proposal.
  5. Guardrail Enforcement & Execution: Robusta validates the proposed action against a strict security policy engine (verifying allowable actions like pod restart, deployment scale, dynamic config update) and applies the fix via the K8s API.

2. Step 1: OpenTelemetry Collector Configuration for Root Cause Telemetry

To provide the LLM agent with precise trace context without overloading the token context window, configure tail-sampling filters in the OpenTelemetry Collector to capture failing spans along with upstream dependency spans.

apiVersion: opentelemetry.io/v1alpha1
kind: OpenTelemetryCollector
metadata:
  name: otel-collector-remediation
  namespace: monitoring
spec:
  config:
    receivers:
      otlp:
        protocols:
          grpc:
            endpoint: 0.0.0.0:4317
          http:
            endpoint: 0.0.0.0:4318

    processors:
      memory_limiter:
        check_interval: 1s
        limit_percentage: 75
        spike_limit_percentage: 15

      tail_sampling:
        decision_wait: 10s
        num_traces: 10000
        expected_new_traces_per_sec: 2000
        policies:
          [
            {
              name: drop-healthchecks,
              type: string_attribute,
              string_attribute: { key: http.target, values: [ "/healthz", "/metrics" ], enabled_regex_matching: false, invert_match: true }
            },
            {
              name: status-code-error,
              type: status_code,
              status_code: { status_codes: [ ERROR ] }
            },
            {
              name: latency-outliers,
              type: latency,
              latency: { threshold_ms: 5000 }
            }
          ]

      batch:
        send_batch_size: 8192
        timeout: 1s

    exporters:
      otlp/tempo:
        endpoint: tempo-distributor.monitoring.svc.cluster.local:4317
        tls:
          insecure: true

      prometheus:
        endpoint: 0.0.0.0:8889
        namespace: otel

    service:
      pipelines:
        traces:
          receivers: [otlp]
          processors: [memory_limiter, tail_sampling, batch]
          exporters: [otlp/tempo]

3. Step 2: Deploying the In-Cluster Local LLM Remediation Agent

Data privacy is critical in enterprise environments. We deploy an in-cluster LLM agent using vLLM to serve DeepSeek-R1-Distill-Qwen-14B or Llama-3-8B-Instruct. The model runs behind a FastAPI wrapper enforcing standard JSON schema outputs.

Deployment Manifest for vLLM Engine

apiVersion: apps/v1
kind: Deployment
metadata:
  name: local-llm-remediation-engine
  namespace: monitoring
spec:
  replicas: 1
  selector:
    matchLabels:
      app: local-llm-engine
  template:
    metadata:
      labels:
        app: local-llm-engine
    spec:
      containers:
      - name: vllm
        image: vllm/vllm-openai:v0.6.3
        args:
        - "--model"
        - "DeepSeek-R1-Distill-Qwen-14B"
        - "--tensor-parallel-size"
        - "1"
        - "--max-model-len"
        - "8192"
        - "--gpu-memory-utilization"
        - "0.90"
        ports:
        - containerPort: 8000
        resources:
          limits:
            nvidia.com/gpu: "1"
            memory: 32Gi
            cpu: "8"
          requests:
            nvidia.com/gpu: "1"
            memory: 16Gi
            cpu: "4"
---
apiVersion: v1
kind: Service
metadata:
  name: local-llm-service
  namespace: monitoring
spec:
  ports:
  - port: 8000
    targetPort: 8000
  selector:
    app: local-llm-engine

4. Step 3: Robusta Custom Action & Automated Remediation Agent

Robusta allows writing custom Python playbooks that extend its reaction engine. The following custom action intercepts a pod crash or memory limit alert, queries OpenTelemetry trace context from Tempo, formats the payload for the local LLM, validates the output schema, and applies the execution.

Robusta Custom Python Action (`llm_remediation.py`)

import json
import requests
from robusta.api import *

class LLMRemediationParams(ActionParams):
    llm_endpoint: str = "http://local-llm-service.monitoring.svc.cluster.local:8000/v1/chat/completions"
    tempo_endpoint: str = "http://tempo-query.monitoring.svc.cluster.local:16686/api/traces"
    auto_execute: bool = True

ALLOWED_ACTIONS = ["restart_pod", "scale_deployment", "patch_environment_variable", "clear_redis_cache"]

@action
def auto_llm_remediate(event: PodEvent, params: LLMRemediationParams):
    pod = event.get_pod()
    if not pod:
        logging.error("No pod found in event trigger.")
        return

    # Extract Kubernetes Logs & Events
    pod_logs = pod.get_logs(tail_lines=100)
    namespace = pod.metadata.namespace
    pod_name = pod.metadata.name

    # Retrieve OTel Trace ID from labels/annotations if available
    trace_id = pod.metadata.annotations.get("opentelemetry.io/last-error-trace-id", None)
    trace_data = {}
    if trace_id:
        try:
            resp = requests.get(f"{params.tempo_endpoint}/{trace_id}", timeout=5)
            if resp.status_code == 200:
                trace_data = resp.json()
        except Exception as e:
            logging.warning(f"Failed to fetch trace context: {str(e)}")

    # Construct Prompt with Schema Enforcement
    system_prompt = (
        "You are an expert Site Reliability Engineer (SRE). Analyze the Kubernetes incident telemetry.
"
        "Select ONE remediation action from this approved list: " + f"{ALLOWED_ACTIONS}
"
        "You MUST respond ONLY with a valid JSON object matching this schema:
"
        "{
"
        '  "root_cause": "string summary",
'
        '  "confidence_score": float (0.0 to 1.0),
'
        '  "action": "allowed_action_name",
'
        '  "target_resource": {"kind": "Deployment|Pod", "name": "string", "namespace": "string"},
'
        '  "parameters": {"key": "value"}
'
        "}"
    )

    user_payload = {
        "pod_name": pod_name,
        "namespace": namespace,
        "logs": pod_logs,
        "trace_context": trace_data
    }

    headers = {"Content-Type": "application/json"}
    body = {
        "model": "DeepSeek-R1-Distill-Qwen-14B",
        "messages": [
            {"role": "system", "content": system_prompt},
            {"role": "user", "content": json.dumps(user_payload)}
        ],
        "temperature": 0.1,
        "response_format": {"type": "json_object"}
    }

    try:
        response = requests.post(params.llm_endpoint, json=body, headers=headers, timeout=30)
        llm_response = response.json()['choices'][0]['message']['content']
        remediation_plan = json.loads(llm_response)

        # Enforce Guardrails
        action = remediation_plan.get("action")
        confidence = remediation_plan.get("confidence_score", 0)

        event.add_enrichment([
            MarkdownBlock(f"### 🤖 LLM Incident Root Cause Analysis
**Cause:** {remediation_plan.get('root_cause')}
**Proposed Action:** `{action}` (Confidence: {confidence})")
        ])

        if confidence >= 0.85 and action in ALLOWED_ACTIONS and params.auto_execute:
            execute_remediation(action, remediation_plan.get("target_resource"), remediation_plan.get("parameters"), event)
        else:
            event.add_enrichment([MarkdownBlock("⚠️ **Action requires manual human approval due to low confidence score or unknown action type.**")])

    except Exception as e:
        logging.error(f"Failed execution of LLM remediation: {str(e)}")

def execute_remediation(action: str, target: dict, params: dict, event: PodEvent):
    if action == "restart_pod":
        RobustaPod.delete_pod(target['name'], target['namespace'])
        event.add_enrichment([MarkdownBlock(f"✅ **Automated Repair Applied:** Deleted pod `{target['name']}` to trigger graceful restart.")])
    elif action == "scale_deployment":
        replicas = int(params.get("replicas", 2))
        RobustaDeployment.scale_deployment(target['name'], target['namespace'], replicas)
        event.add_enrichment([MarkdownBlock(f"✅ **Automated Repair Applied:** Scaled deployment `{target['name']}` to `{replicas}` replicas.")])

5. Step 4: Configuring Robusta Playbooks and Security RBAC

To bind this custom action to incoming cluster alerts, define a custom playbook in the Robusta Helm configuration (`generated_values.yaml`).

custom_playbooks:
  - name: "llm-automated-remediation"
    triggers:
      - on_pod_crash_loop: {}
      - on_prometheus_alert:
          alert_name: "KubeContainerWaiting"
      - on_prometheus_alert:
          alert_name: "HighHttp5xxErrorRate"
    actions:
      - auto_llm_remediate:
          llm_endpoint: "http://local-llm-service.monitoring.svc.cluster.local:8000/v1/chat/completions"
          tempo_endpoint: "http://tempo-query.monitoring.svc.cluster.local:16686/api/traces"
          auto_execute: true

Least-Privilege Security RBAC

Ensure the Robusta ServiceAccount operates under strict, explicit role boundaries. Never grant cluster-admin access to an automated remediation runner.

apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
  name: robusta-remediation-runner
rules:
- apiGroups: [""]
  resources: ["pods", "pods/log", "events"]
  verbs: ["get", "list", "watch", "delete"]
- apiGroups: ["apps"]
  resources: ["deployments", "statefulsets"]
  verbs: ["get", "list", "update", "patch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
  name: robusta-remediation-runner-binding
subjects:
- kind: ServiceAccount
  name: robusta-forwarder-service-account
  namespace: monitoring
roleRef:
  kind: ClusterRole
  name: robusta-remediation-runner
  apiGroup: rbac.authorization.k8s.io

6. Enterprise Guardrails & Production Verification

Security / Reliability VectorFailure ModeEnterprise Mitigation Strategy
LLM HallucinationModel invents non-existent K8s CLI commands or destructive API flags.Enforce strict JSON schema parsing and whitelist permissible action function endpoints in the Python wrapper.
Flapping / Remediation LoopsPod repeatedly crashes, causing cyclic automated restarts.Implement exponential backoff and rate-limiting triggers inside Robusta (max 2 repairs per resource per hour).
Data Privacy BreachSending production trace payloads containing PII to external APIs.Host the model locally within the K8s cluster (vLLM / Ollama) behind strict network policies; no egress traffic.
Unauthorized EscalationLLM targets unauthorized core system namespaces.Kubernetes RBAC restrictions prevent modifying resources in `kube-system` or `monitoring`.

Verification Procedure

Verify the self-healing workflow by injecting a controlled memory-leak failure into a test microservice:

# 1. Trigger a fault injection pod
kubectl run fault-injector --image=busybox --namespace=default -- /bin/sh -c "memhog 500M"

# 2. Watch Robusta remediation logs
kubectl logs -n monitoring -l app=robusta-runner -f

# 3. Check Slack or Alertmanager output for structured LLM enrichment
# Expected output: Root Cause determined -> Graceful target restart applied.

Conclusion

By coupling OpenTelemetry trace context with Robusta automation and an in-cluster Local LLM agent, SRE teams can move beyond passive monitoring to proactive, self-healing platforms. The model provides intelligent, context-aware diagnostics based on real-time distributed traces, while deterministic guardrails and strict RBAC ensure safe, enterprise-compliant execution.