Hạ Tầng AI Tự Phục Hồi: Tự Động Khắc Phục Sự Cố Bằng OpenTelemetry, Robusta và LLM Cục Bộ
Hạ Tầng AI Tự Phục Hồi: Tự Động Khắc Phục Sự Cố Bằng OpenTelemetry Spans, Robusta và K8s Local LLM Agents
Các cụm Kubernetes doanh nghiệp hiện đại tạo ra lượng dữ liệu đo đạc (telemetry) khổng lồ. Tuy nhiên, Thời gian Phát hiện Trung bình (MTTD) và Thời gian Khắc phục Trung bình (MTTR) bị nghẽn nghiêm trọng bởi năng lực xử lý thủ công của con người. Khi sự cố xảy ra—như cạn kiệt connection pool dây chuyền hoặc rò rỉ bộ nhớ trong microservice—SRE phải tương quan vết OpenTelemetry, lấy log pod, phân tích sự kiện Kubernetes và thực hiện thao tác sửa lỗi bằng tay.
Bài viết này hướng dẫn chi tiết kiến trúc cấp doanh nghiệp cho Hạ tầng Kubernetes Tự Phục Hồi (Self-Healing). Bằng cách kết hợp ngữ cảnh OpenTelemetry (OTel) trace, tự động hóa sự cố Robusta.dev, và Mô hình Ngôn ngữ Lớn cục bộ (Local LLM Agent) (như DeepSeek-R1 hoặc Llama-3 chạy trên vLLM ngay trong cluster), chúng ta xây dựng đường ống khắc phục tự động khép kín: tự phân tích nguyên nhân gốc và thực thi hành động sửa chữa an toàn mà không đưa dữ liệu nhạy cảm ra ngoài vây bọc bảo mật.
2. Tổng Quan Kiến Trúc & Luồng Hoạt ĐộngKiến trúc tự phục hồi tách biệt giữa việc thu thập vết telemetry, làm giàu cảnh báo, suy luận thông minh và rào chắn thực thi an toàn.
+-----------------------------------------------------------------------------------+
| CỤM KUBERNETES |
| |
| +--------------------+ OTLP +---------------------------------------+ |
| | Application Pods | -------------> | OpenTelemetry Collector | |
| | (OTel Instrumented)| | (Tail-sampling & Span Metrics) | |
| +--------------------+ +---------------------------------------+ |
| | | |
| Request Kích hoạt Alert |
| Thất bại | |
| v v |
| +--------------------+ +-------------------------------+ |
| | K8s Events / Logs | <--------------------- | Prometheus Alertmanager | |
| +--------------------+ +-------------------------------+ |
| | |
| Alert Payload |
| v |
| +-------------------------------+ |
| | Robusta Operator | |
| | (Enrichment Engine) | |
| +-------------------------------+ |
| | |
| Ngữ cảnh Span + Log |
| v |
| +-------------------------------+ |
| | Local LLM Agent | |
| | (vLLM / Ollama + API Nội bộ) | |
| +-------------------------------+ |
| | |
| JSON Hành động Chuẩn |
| v |
| +-------------------------------+ |
| | Robusta Remediation Runner | |
| | (RBAC & Guardrail Enforcer) | |
| +-------------------------------+ |
| | |
| Thực thi API Patch |
| v |
| +-------------------------------+ |
| | Kubernetes API Server | |
| +-------------------------------+ |
+-----------------------------------------------------------------------------------+Vòng Đời Của Luồng Tự Khắc Phục
- Thu Thập Telemetry & Trích Xuất Ngữ Cảnh: Pod ứng dụng phát ra span qua OTLP. OTel Collector lọc ra các error span (như HTTP 5xx, DB timeout) và đính kèm `trace_id`, `span_id`.
- Kích Hoạt Làm Giàu Cảnh Báo: Metric từ Prometheus hoặc OTel Collector đẩy alert tới Prometheus Alertmanager, sau đó chuyển tiếp tới engine Robusta.
- Thu Thập Ngữ Cảnh Sâu: Robusta truy vấn Tempo API để lấy đầy đủ trace span dẫn tới sự cố, đồng thời thu thập 100 dòng log cuối của pod và các sự kiện K8s liên quan.
- Phân Tích Bởi LLM Cục Bộ: Robusta gửi ngữ cảnh sự cố cấu trúc tới dịch vụ inference LLM nội bộ (vLLM chạy model quantized). Model trả về bản đề xuất khắc phục theo cấu trúc JSON nghiêm ngặt.
- Kiểm Tra Rào Chắn & Thực Thi: Robusta xác minh hành động đối chiếu với chính sách bảo mật (chỉ cho phép hành động như restart pod, scale deployment, patch config) và gửi lệnh tới K8s API.
2. Bước 1: Cấu Hình OpenTelemetry Collector Để Lấy Telemetry Nguyên Nhân Gốc
Để cung cấp ngữ cảnh vết chính xác cho LLM agent mà không làm tràn context window, cấu hình bộ lọc tail-sampling trong OpenTelemetry Collector nhằm bắt các span lỗi cùng các span phụ thuộc phía trước.
apiVersion: opentelemetry.io/v1alpha1
kind: OpenTelemetryCollector
metadata:
name: otel-collector-remediation
namespace: monitoring
spec:
config:
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
processors:
memory_limiter:
check_interval: 1s
limit_percentage: 75
spike_limit_percentage: 15
tail_sampling:
decision_wait: 10s
num_traces: 10000
expected_new_traces_per_sec: 2000
policies:
[
{
name: drop-healthchecks,
type: string_attribute,
string_attribute: { key: http.target, values: [ "/healthz", "/metrics" ], enabled_regex_matching: false, invert_match: true }
},
{
name: status-code-error,
type: status_code,
status_code: { status_codes: [ ERROR ] }
},
{
name: latency-outliers,
type: latency,
latency: { threshold_ms: 5000 }
}
]
batch:
send_batch_size: 8192
timeout: 1s
exporters:
otlp/tempo:
endpoint: tempo-distributor.monitoring.svc.cluster.local:4317
tls:
insecure: true
prometheus:
endpoint: 0.0.0.0:8889
namespace: otel
service:
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, tail_sampling, batch]
exporters: [otlp/tempo]3. Bước 2: Triển Khai Engine Local LLM Trong Cụm
Quyền riêng tư dữ liệu là yếu tố sống còn trong môi trường doanh nghiệp. Chúng ta triển khai LLM agent ngay trong cluster sử dụng vLLM để phục vụ model DeepSeek-R1-Distill-Qwen-14B hoặc Llama-3-8B-Instruct. Engine này nằm sau một FastAPI wrapper để bắt buộc định dạng đầu ra đúng JSON schema.
File Triển Khai vLLM Engine
apiVersion: apps/v1
kind: Deployment
metadata:
name: local-llm-remediation-engine
namespace: monitoring
spec:
replicas: 1
selector:
matchLabels:
app: local-llm-engine
template:
metadata:
labels:
app: local-llm-engine
spec:
containers:
- name: vllm
image: vllm/vllm-openai:v0.6.3
args:
- "--model"
- "DeepSeek-R1-Distill-Qwen-14B"
- "--tensor-parallel-size"
- "1"
- "--max-model-len"
- "8192"
- "--gpu-memory-utilization"
- "0.90"
ports:
- containerPort: 8000
resources:
limits:
nvidia.com/gpu: "1"
memory: 32Gi
cpu: "8"
requests:
nvidia.com/gpu: "1"
memory: 16Gi
cpu: "4"
---
apiVersion: v1
kind: Service
metadata:
name: local-llm-service
namespace: monitoring
spec:
ports:
- port: 8000
targetPort: 8000
selector:
app: local-llm-engine4. Bước 3: Robusta Custom Action & Python Remediation Agent
Robusta cho phép viết custom Python playbook để mở rộng khả năng phản ứng. Mã nguồn Python dưới đây chặn alert Pod crash/OOM, lấy dữ liệu OTel trace từ Tempo, định dạng prompt cho Local LLM, kiểm tra schema và thực thi giải pháp.
Custom Action Python (`llm_remediation.py`)
import json
import requests
from robusta.api import *
class LLMRemediationParams(ActionParams):
llm_endpoint: str = "http://local-llm-service.monitoring.svc.cluster.local:8000/v1/chat/completions"
tempo_endpoint: str = "http://tempo-query.monitoring.svc.cluster.local:16686/api/traces"
auto_execute: bool = True
ALLOWED_ACTIONS = ["restart_pod", "scale_deployment", "patch_environment_variable", "clear_redis_cache"]
@action
def auto_llm_remediate(event: PodEvent, params: LLMRemediationParams):
pod = event.get_pod()
if not pod:
logging.error("Không tìm thấy pod trong sự kiện kích hoạt.")
return
# Trích xuất Log và Event của Kubernetes
pod_logs = pod.get_logs(tail_lines=100)
namespace = pod.metadata.namespace
pod_name = pod.metadata.name
# Lấy Trace ID từ Annotation nếu có
trace_id = pod.metadata.annotations.get("opentelemetry.io/last-error-trace-id", None)
trace_data = {}
if trace_id:
try:
resp = requests.get(f"{params.tempo_endpoint}/{trace_id}", timeout=5)
if resp.status_code == 200:
trace_data = resp.json()
except Exception as e:
logging.warning(f"Không thể lấy ngữ cảnh trace: {str(e)}")
# Xây dựng Prompt bắt buộc định dạng JSON
system_prompt = (
"Bạn là một Kỹ sư SRE chuyên nghiệp. Hãy phân tích dữ liệu telemetry sự cố Kubernetes.
"
"Chỉ chọn MỘT hành động khắc phục từ danh sách được phép sau: " + f"{ALLOWED_ACTIONS}
"
"Bạn BẮT BUỘC chỉ trả về duy nhất một JSON object hợp lệ theo schema:
"
"{
"
' "root_cause": "tóm tắt nguyên nhân",
'
' "confidence_score": float (0.0 đến 1.0),
'
' "action": "tên_hành_động_được_phép",
'
' "target_resource": {"kind": "Deployment|Pod", "name": "string", "namespace": "string"},
'
' "parameters": {"key": "value"}
'
"}"
)
user_payload = {
"pod_name": pod_name,
"namespace": namespace,
"logs": pod_logs,
"trace_context": trace_data
}
headers = {"Content-Type": "application/json"}
body = {
"model": "DeepSeek-R1-Distill-Qwen-14B",
"messages": [
{"role": "system", "content": system_prompt},
{"role": "user", "content": json.dumps(user_payload)}
],
"temperature": 0.1,
"response_format": {"type": "json_object"}
}
try:
response = requests.post(params.llm_endpoint, json=body, headers=headers, timeout=30)
llm_response = response.json()['choices'][0]['message']['content']
remediation_plan = json.loads(llm_response)
# Kiểm tra Guardrail
action = remediation_plan.get("action")
confidence = remediation_plan.get("confidence_score", 0)
event.add_enrichment([
MarkdownBlock(f"### 🤖 LLM Phân Tích Nguyên Nhân Gốc Sự Cố
**Nguyên nhân:** {remediation_plan.get('root_cause')}
**Đề xuất:** `{action}` (Độ tin cậy: {confidence})")
])
if confidence >= 0.85 and action in ALLOWED_ACTIONS and params.auto_execute:
execute_remediation(action, remediation_plan.get("target_resource"), remediation_plan.get("parameters"), event)
else:
event.add_enrichment([MarkdownBlock("⚠️ **Hành động yêu cầu con người phê duyệt do điểm tin cậy thấp hoặc hành động không nằm trong danh mục.**")])
except Exception as e:
logging.error(f"Thực thi khắc phục sự cố thất bại: {str(e)}")
def execute_remediation(action: str, target: dict, params: dict, event: PodEvent):
if action == "restart_pod":
RobustaPod.delete_pod(target['name'], target['namespace'])
event.add_enrichment([MarkdownBlock(f"✅ **Đã tự động khắc phục:** Xóa pod `{target['name']}` để khởi động lại an toàn.")])
elif action == "scale_deployment":
replicas = int(params.get("replicas", 2))
RobustaDeployment.scale_deployment(target['name'], target['namespace'], replicas)
event.add_enrichment([MarkdownBlock(f"✅ **Đã tự động khắc phục:** Scaled deployment `{target['name']}` lên `{replicas}` replicas.")])5. Bước 4: Cấu Hình Playbook Robusta Và Phân Quyền RBAC
Để liên kết custom action này với các alert của cluster, khai báo playbook trong cấu hình Robusta Helm (`generated_values.yaml`).
custom_playbooks:
- name: "llm-automated-remediation"
triggers:
- on_pod_crash_loop: {}
- on_prometheus_alert:
alert_name: "KubeContainerWaiting"
- on_prometheus_alert:
alert_name: "HighHttp5xxErrorRate"
actions:
- auto_llm_remediate:
llm_endpoint: "http://local-llm-service.monitoring.svc.cluster.local:8000/v1/chat/completions"
tempo_endpoint: "http://tempo-query.monitoring.svc.cluster.local:16686/api/traces"
auto_execute: trueBảo Mật Quyền Tối Thiểu (RBAC)
Gán quyền cho Robusta ServiceAccount theo cơ chế tối thiểu. Tuyệt đối không cấp quyền `cluster-admin` cho bot tự động sửa lỗi.
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: robusta-remediation-runner
rules:
- apiGroups: [""]
resources: ["pods", "pods/log", "events"]
verbs: ["get", "list", "watch", "delete"]
- apiGroups: ["apps"]
resources: ["deployments", "statefulsets"]
verbs: ["get", "list", "update", "patch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: robusta-remediation-runner-binding
subjects:
- kind: ServiceAccount
name: robusta-forwarder-service-account
namespace: monitoring
roleRef:
kind: ClusterRole
name: robusta-remediation-runner
apiGroup: rbac.authorization.k8s.io6. Chiến Lược Rào Chắn Doanh Nghiệp (Guardrails) & Kiểm Thử
| Rủi Ro Bảo Mật / Tin Cậy | Kịch Bản Thất Bại | Chiến Lược Khắc Phục Doanh Nghiệp |
|---|---|---|
| LLM Hallucination (Ảo giác) | Model tự bịa ra lệnh K8s CLI hoặc tham số phá hoại. | Bắt buộc parse JSON schema và sử dụng danh sách whitelist hàm được phép trong Python wrapper. |
| Lặp Khắc Phục (Flapping) | Pod crash liên tục gây ra vòng lặp restart tự động. | Thiết lập rate-limit và exponential backoff trong Robusta (tối đa 2 lần sửa / resource / giờ). |
| Rò Rỉ Dữ Liệu | Gửi trace payload có chứa PII ra các API bên ngoài. | Chạy LLM nội bộ trong cluster (vLLM / Ollama) chặn hoàn toàn egress internet. |
| Lleo Thang Quyền Bất Hợp Pháp | LLM can thiệp vào namespace hệ thống cốt lõi. | Chính sách K8s RBAC ngăn chặn sửa đổi resource trong `kube-system` hoặc `monitoring`. |
Quy Trình Kiểm Thử Thực Tế
Kiểm tra luồng tự phục hồi bằng cách chủ động giả lập sự cố rò rỉ bộ nhớ vào microservice thử nghiệm:
# 1. Chạy pod giả lập sự cố
kubectl run fault-injector --image=busybox --namespace=default -- /bin/sh -c "memhog 500M"
# 2. Theo dõi log của Robusta remediation runner
kubectl logs -n monitoring -l app=robusta-runner -f
# 3. Kiểm tra Slack hoặc Alertmanager để xem kết quả phân tích enrichment từ LLM
# Kết quả kỳ vọng: LLM xác định đúng nguyên nhân -> Thực thi khởi động lại target pod thành công.Kết Luận
Bằng cách kết hợp ngữ cảnh vết OpenTelemetry với engine tự động hóa Robusta và Local LLM agent trong cụm, đội ngũ SRE có thể chuyển dịch từ giám sát thụ động sang hạ tầng tự chữa lành chủ động. LLM mang lại khả năng chẩn đoán thông minh, chính xác dựa trên trace phân tán thời gian thực, trong khi các rào chắn quy tắc và RBAC đảm bảo quá trình thực thi luôn an toàn và tuân thủ tiêu chuẩn doanh nghiệp.
