Self-Healing AI Infrastructure: Automated Incident Remediation with OpenTelemetry, Robusta & Local LLMs
Self-Healing AI Infrastructure: Automated Incident Remediation with OpenTelemetry Spans, Robusta, and K8s Local LLM Agents
Modern enterprise Kubernetes clusters generate immense telemetry volume. However, traditional Mean Time To Detection (MTTD) and Mean Time To Resolution (MTTR) are severely bottlenecked by human cognitive bandwidth. When an incident occurs—such as a cascading connection pool exhaustion or a memory leak in a microservice—SREs must manually correlate OpenTelemetry traces, pull pod logs, analyze Kubernetes events, and execute patch maneuvers.
This guide presents an enterprise-grade architectural blueprint for Self-Healing Kubernetes Infrastructure. By combining OpenTelemetry (OTel) trace context, Robusta.dev incident automation, and an in-cluster Local Large Language Model (LLM) agent (e.g., DeepSeek-R1 or Llama-3 running on vLLM), we build a closed-loop remediation pipeline that autonomously analyzes root causes and executes deterministic, guarded repairs without sending sensitive telemetry outside your security perimeter.
1. Architectural Overview & Workflow
The self-healing architecture decouples trace telemetry generation, alert enrichment, intelligent reasoning, and safe execution guardrails.
+-----------------------------------------------------------------------------------+
| KUBERNETES CLUSTER |
| |
| +--------------------+ OTLP +---------------------------------------+ |
| | Application Pods | -------------> | OpenTelemetry Collector | |
| | (OTel Instrumented)| | (Tail-sampling & Span Metrics) | |
| +--------------------+ +---------------------------------------+ |
| | | |
| Failing Triggers Alert |
| Request | |
| v v |
| +--------------------+ +-------------------------------+ |
| | K8s Events / Logs | <--------------------- | Prometheus Alertmanager | |
| +--------------------+ +-------------------------------+ |
| | |
| Alert Payload |
| v |
| +-------------------------------+ |
| | Robusta Operator | |
| | (Enrichment Engine) | |
| +-------------------------------+ |
| | |
| Span Context + Logs |
| v |
| +-------------------------------+ |
| | In-Cluster LLM Remediation | |
| | Agent (vLLM / Ollama + API) | |
| +-------------------------------+ |
| | |
| Strict JSON Action |
| v |
| +-------------------------------+ |
| | Robusta Remediation Runner | |
| | (RBAC & Guardrail Enforcer) | |
| +-------------------------------+ |
| | |
| Executes API Patch |
| v |
| +-------------------------------+ |
| | Kubernetes API Server | |
| +-------------------------------+ |
+-----------------------------------------------------------------------------------+Remediation Flow Lifecycle
- Telemetry Capture & Context Extraction: Application pods emit spans via OTLP. The OTel Collector identifies error spans (e.g., HTTP 5xx, DB timeout) and attaches contextual IDs (`trace_id`, `span_id`).
- Enrichment Trigger: Prometheus or OTel Collector metrics fire an alert to Prometheus Alertmanager, which forwards it to the Robusta engine.
- Deep Enrichment: Robusta queries the Jaeger/Tempo API for full trace spans leading up to the crash and collects the last 100 lines of pod stdout/stderr alongside Kubernetes events.
- Local LLM Analysis: Robusta sends the structured incident context to a local, isolated LLM inference service (vLLM running a quantized model). The model outputs a strictly structured JSON remediation proposal.
- Guardrail Enforcement & Execution: Robusta validates the proposed action against a strict security policy engine (verifying allowable actions like pod restart, deployment scale, dynamic config update) and applies the fix via the K8s API.
2. Step 1: OpenTelemetry Collector Configuration for Root Cause Telemetry
To provide the LLM agent with precise trace context without overloading the token context window, configure tail-sampling filters in the OpenTelemetry Collector to capture failing spans along with upstream dependency spans.
apiVersion: opentelemetry.io/v1alpha1
kind: OpenTelemetryCollector
metadata:
name: otel-collector-remediation
namespace: monitoring
spec:
config:
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
processors:
memory_limiter:
check_interval: 1s
limit_percentage: 75
spike_limit_percentage: 15
tail_sampling:
decision_wait: 10s
num_traces: 10000
expected_new_traces_per_sec: 2000
policies:
[
{
name: drop-healthchecks,
type: string_attribute,
string_attribute: { key: http.target, values: [ "/healthz", "/metrics" ], enabled_regex_matching: false, invert_match: true }
},
{
name: status-code-error,
type: status_code,
status_code: { status_codes: [ ERROR ] }
},
{
name: latency-outliers,
type: latency,
latency: { threshold_ms: 5000 }
}
]
batch:
send_batch_size: 8192
timeout: 1s
exporters:
otlp/tempo:
endpoint: tempo-distributor.monitoring.svc.cluster.local:4317
tls:
insecure: true
prometheus:
endpoint: 0.0.0.0:8889
namespace: otel
service:
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, tail_sampling, batch]
exporters: [otlp/tempo]3. Step 2: Deploying the In-Cluster Local LLM Remediation Agent
Data privacy is critical in enterprise environments. We deploy an in-cluster LLM agent using vLLM to serve DeepSeek-R1-Distill-Qwen-14B or Llama-3-8B-Instruct. The model runs behind a FastAPI wrapper enforcing standard JSON schema outputs.
Deployment Manifest for vLLM Engine
apiVersion: apps/v1
kind: Deployment
metadata:
name: local-llm-remediation-engine
namespace: monitoring
spec:
replicas: 1
selector:
matchLabels:
app: local-llm-engine
template:
metadata:
labels:
app: local-llm-engine
spec:
containers:
- name: vllm
image: vllm/vllm-openai:v0.6.3
args:
- "--model"
- "DeepSeek-R1-Distill-Qwen-14B"
- "--tensor-parallel-size"
- "1"
- "--max-model-len"
- "8192"
- "--gpu-memory-utilization"
- "0.90"
ports:
- containerPort: 8000
resources:
limits:
nvidia.com/gpu: "1"
memory: 32Gi
cpu: "8"
requests:
nvidia.com/gpu: "1"
memory: 16Gi
cpu: "4"
---
apiVersion: v1
kind: Service
metadata:
name: local-llm-service
namespace: monitoring
spec:
ports:
- port: 8000
targetPort: 8000
selector:
app: local-llm-engine4. Step 3: Robusta Custom Action & Automated Remediation Agent
Robusta allows writing custom Python playbooks that extend its reaction engine. The following custom action intercepts a pod crash or memory limit alert, queries OpenTelemetry trace context from Tempo, formats the payload for the local LLM, validates the output schema, and applies the execution.
Robusta Custom Python Action (`llm_remediation.py`)
import json
import requests
from robusta.api import *
class LLMRemediationParams(ActionParams):
llm_endpoint: str = "http://local-llm-service.monitoring.svc.cluster.local:8000/v1/chat/completions"
tempo_endpoint: str = "http://tempo-query.monitoring.svc.cluster.local:16686/api/traces"
auto_execute: bool = True
ALLOWED_ACTIONS = ["restart_pod", "scale_deployment", "patch_environment_variable", "clear_redis_cache"]
@action
def auto_llm_remediate(event: PodEvent, params: LLMRemediationParams):
pod = event.get_pod()
if not pod:
logging.error("No pod found in event trigger.")
return
# Extract Kubernetes Logs & Events
pod_logs = pod.get_logs(tail_lines=100)
namespace = pod.metadata.namespace
pod_name = pod.metadata.name
# Retrieve OTel Trace ID from labels/annotations if available
trace_id = pod.metadata.annotations.get("opentelemetry.io/last-error-trace-id", None)
trace_data = {}
if trace_id:
try:
resp = requests.get(f"{params.tempo_endpoint}/{trace_id}", timeout=5)
if resp.status_code == 200:
trace_data = resp.json()
except Exception as e:
logging.warning(f"Failed to fetch trace context: {str(e)}")
# Construct Prompt with Schema Enforcement
system_prompt = (
"You are an expert Site Reliability Engineer (SRE). Analyze the Kubernetes incident telemetry.
"
"Select ONE remediation action from this approved list: " + f"{ALLOWED_ACTIONS}
"
"You MUST respond ONLY with a valid JSON object matching this schema:
"
"{
"
' "root_cause": "string summary",
'
' "confidence_score": float (0.0 to 1.0),
'
' "action": "allowed_action_name",
'
' "target_resource": {"kind": "Deployment|Pod", "name": "string", "namespace": "string"},
'
' "parameters": {"key": "value"}
'
"}"
)
user_payload = {
"pod_name": pod_name,
"namespace": namespace,
"logs": pod_logs,
"trace_context": trace_data
}
headers = {"Content-Type": "application/json"}
body = {
"model": "DeepSeek-R1-Distill-Qwen-14B",
"messages": [
{"role": "system", "content": system_prompt},
{"role": "user", "content": json.dumps(user_payload)}
],
"temperature": 0.1,
"response_format": {"type": "json_object"}
}
try:
response = requests.post(params.llm_endpoint, json=body, headers=headers, timeout=30)
llm_response = response.json()['choices'][0]['message']['content']
remediation_plan = json.loads(llm_response)
# Enforce Guardrails
action = remediation_plan.get("action")
confidence = remediation_plan.get("confidence_score", 0)
event.add_enrichment([
MarkdownBlock(f"### 🤖 LLM Incident Root Cause Analysis
**Cause:** {remediation_plan.get('root_cause')}
**Proposed Action:** `{action}` (Confidence: {confidence})")
])
if confidence >= 0.85 and action in ALLOWED_ACTIONS and params.auto_execute:
execute_remediation(action, remediation_plan.get("target_resource"), remediation_plan.get("parameters"), event)
else:
event.add_enrichment([MarkdownBlock("⚠️ **Action requires manual human approval due to low confidence score or unknown action type.**")])
except Exception as e:
logging.error(f"Failed execution of LLM remediation: {str(e)}")
def execute_remediation(action: str, target: dict, params: dict, event: PodEvent):
if action == "restart_pod":
RobustaPod.delete_pod(target['name'], target['namespace'])
event.add_enrichment([MarkdownBlock(f"✅ **Automated Repair Applied:** Deleted pod `{target['name']}` to trigger graceful restart.")])
elif action == "scale_deployment":
replicas = int(params.get("replicas", 2))
RobustaDeployment.scale_deployment(target['name'], target['namespace'], replicas)
event.add_enrichment([MarkdownBlock(f"✅ **Automated Repair Applied:** Scaled deployment `{target['name']}` to `{replicas}` replicas.")])5. Step 4: Configuring Robusta Playbooks and Security RBAC
To bind this custom action to incoming cluster alerts, define a custom playbook in the Robusta Helm configuration (`generated_values.yaml`).
custom_playbooks:
- name: "llm-automated-remediation"
triggers:
- on_pod_crash_loop: {}
- on_prometheus_alert:
alert_name: "KubeContainerWaiting"
- on_prometheus_alert:
alert_name: "HighHttp5xxErrorRate"
actions:
- auto_llm_remediate:
llm_endpoint: "http://local-llm-service.monitoring.svc.cluster.local:8000/v1/chat/completions"
tempo_endpoint: "http://tempo-query.monitoring.svc.cluster.local:16686/api/traces"
auto_execute: trueLeast-Privilege Security RBAC
Ensure the Robusta ServiceAccount operates under strict, explicit role boundaries. Never grant cluster-admin access to an automated remediation runner.
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: robusta-remediation-runner
rules:
- apiGroups: [""]
resources: ["pods", "pods/log", "events"]
verbs: ["get", "list", "watch", "delete"]
- apiGroups: ["apps"]
resources: ["deployments", "statefulsets"]
verbs: ["get", "list", "update", "patch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: robusta-remediation-runner-binding
subjects:
- kind: ServiceAccount
name: robusta-forwarder-service-account
namespace: monitoring
roleRef:
kind: ClusterRole
name: robusta-remediation-runner
apiGroup: rbac.authorization.k8s.io6. Enterprise Guardrails & Production Verification
| Security / Reliability Vector | Failure Mode | Enterprise Mitigation Strategy |
|---|---|---|
| LLM Hallucination | Model invents non-existent K8s CLI commands or destructive API flags. | Enforce strict JSON schema parsing and whitelist permissible action function endpoints in the Python wrapper. |
| Flapping / Remediation Loops | Pod repeatedly crashes, causing cyclic automated restarts. | Implement exponential backoff and rate-limiting triggers inside Robusta (max 2 repairs per resource per hour). |
| Data Privacy Breach | Sending production trace payloads containing PII to external APIs. | Host the model locally within the K8s cluster (vLLM / Ollama) behind strict network policies; no egress traffic. |
| Unauthorized Escalation | LLM targets unauthorized core system namespaces. | Kubernetes RBAC restrictions prevent modifying resources in `kube-system` or `monitoring`. |
Verification Procedure
Verify the self-healing workflow by injecting a controlled memory-leak failure into a test microservice:
# 1. Trigger a fault injection pod
kubectl run fault-injector --image=busybox --namespace=default -- /bin/sh -c "memhog 500M"
# 2. Watch Robusta remediation logs
kubectl logs -n monitoring -l app=robusta-runner -f
# 3. Check Slack or Alertmanager output for structured LLM enrichment
# Expected output: Root Cause determined -> Graceful target restart applied.Conclusion
By coupling OpenTelemetry trace context with Robusta automation and an in-cluster Local LLM agent, SRE teams can move beyond passive monitoring to proactive, self-healing platforms. The model provides intelligent, context-aware diagnostics based on real-time distributed traces, while deterministic guardrails and strict RBAC ensure safe, enterprise-compliant execution.
