Quay lại danh sách
Tin tức công nghệ

Hạ Tầng AI Tự Chữa Lành: Tự Động Khắc Phục Sự Cố Với OpenTelemetry, Robusta Và Local LLM Agent

22 tháng 8, 2026

Hạ Tầng AI Tự Chữa Lành: Tự Động Khắc Phục Sự Cố Với OpenTelemetry, Robusta Và Local LLM Agent

Vận hành hạ tầng AI quy mô lớn—như các cụm inference vLLM phân tán, pipeline fine-tuning, và cơ sở dữ liệu vector throughput cao—đặt ra những thách thức vận hành hoàn toàn khác biệt so với các nền tảng microservice truyền thống. Các cụm node GPU mật độ cao thường gặp phải các lỗi rất khó chẩn đoán: treo cứng do CUDA Out-Of-Memory (OOM), quá thời gian đồng bộ hóa giao tiếp NCCL, tranh chấp KV-cache, phân mảnh bộ nhớ GPU và lỗi PCI bus tức thời. Hệ thống cảnh báo tĩnh truyền thống thường thất bại trong các môi trường này, gây ra tình trạng nhiễu cảnh báo (alert fatigue) trong khi các pipeline AI quan trọng bị đình trệ.

Bài viết này hướng dẫn kiến trúc chi tiết xây dựng nền tảng Hạ tầng AI tự chữa lành (Self-Healing AI Infrastructure) khép kín. Bằng cách kết hợp dữ liệu vết phân tán (distributed tracing) chi tiết từ OpenTelemetry (OTel), khả năng điều phối theo sự kiện từ Robusta, cùng khả năng suy luận ngữ cảnh từ một Local LLM Agent đặt trong cụm K8s (chạy bằng vLLM và DeepSeek-Coder/Mistral bảo mật dữ liệu), các kỹ sư platform có thể tự động hóa việc chẩn đoán và khắc phục sự cố theo thời gian thực mà không cần sự can thiệp của con người hoặc tiết lộ telemetry ra các API bên ngoài.

Tổng Quan Kiến Trúc & Luồng Điều Khiển

Nền tảng thiết lập một vòng phản hồi theo sự kiện, biến các lỗi hệ thống phi cấu trúc thành các pipeline thực thi xác định:

+-----------------------------------------------------------------------------------+
|                            Kubernetes AI Workload Cluster                         |
|                                                                                   |
|  +----------------------+      +----------------------+     +------------------+  |
|  | vLLM / Fine-Tuning   |      | NVIDIA DCGM Exporter |     | Vector DB Pods   |  |
|  | (OTel Instrumentation|      | (GPU Metrics)        |     | (Qdrant/Milvus)  |  |
|  +----------+-----------+      +----------+-----------+     +--------+---------+  |
|             | Spans & Metrics             | Metric Streams           |            |
|             +-----------------+-----------+                          |            |
|                               v                                      v            |
|                     +-------------------+                                         |
|                     | OpenTelemetry     |                                         |
|                     | Collector Daemon  |                                         |
|                     +---------+---------+                                         |
|                               | OTLP Traces & PromMetrics                         |
|                               v                                                   |
|                     +-------------------+                                         |
|                     | Prometheus &      |                                         |
|                     | Alertmanager      |                                         |
|                     +---------+---------+                                         |
|                               | Webhook Alerts                                    |
|                               v                                                   |
|                     +-------------------+                                         |
|                     | Robusta Engine    |                                         |
|                     | (Automation Hub)  |                                         |
|                     +---------+---------+                                         |
|                               | Rich Context JSON Payload                         |
|                               v                                                   |
|                     +-------------------+        +-----------------------------+  |
|                     | Local LLM Agent   |------->| Local LLM Engine (vLLM)     |  |
|                     | (ReAct Controller)|<-------| (DeepSeek-Coder-33B)        |  |
|                     +---------+---------+        +-----------------------------+  |
|                               | Function Call / Remediation Intent                |
|                               v                                                   |
|                     +-------------------+                                         |
|                     | Kubernetes API /  |--> Resets GPU / Adjusts KV-Cache /      |
|                     | NVML Operator     |    Evicts Pod / Taints Bad Node         |
|                     +-------------------+                                         |
+-----------------------------------------------------------------------------------+

Vòng Đời Vận Hành Tự Chữa Lành

  1. Phát sinh Telemetry: Ứng dụng AI phát ra các span OpenTelemetry phân tán chi tiết về độ trễ tạo token (Time-To-First-Token - TTFT), phân bổ KV-cache, và các hàm NCCL. NVIDIA DCGM xuất ra telemetry thiết bị GPU mức thấp.
  2. Phát hiện Bất thường: OTel Collector xử lý các trace span; Prometheus Alertmanager phát hiện vi phạm ngưỡng hoặc dấu hiệu bất thường dựa trên trace (ví dụ: độ trễ đuôi tăng cao kết hợp với phân mảnh bộ nhớ GPU).
  3. Bổ sung Ngữ cảnh & Điều hướng: Robusta chặn các sự kiện từ Alertmanager, gom nhóm cây trace đầy đủ, log gần nhất của pod (`nvidia-smi`, kernel log), và gửi payload JSON đã cấu trúc sang AI Remediation Agent nội bộ.
  4. Suy luận LLM Agent & Kiểm soát: Local LLM Agent phân tích sự cố dựa trên JSON schema nghiêm ngặt và playbook vận hành, xuất ra hành động khắc phục chính xác.
  5. Tự động Thực thi: Agent thực thi patch trực tiếp lên Kubernetes API, kích hoạt reset GPU mức host qua NVML daemonset, hoặc cập nhật lại cấu hình động của deployment.

Bước 1: Cấu Hình OpenTelemetry Cho Pipeline AI

Các metric HTTP thông thường không thể phản ánh trạng thái runtime của GPU. Chúng ta phải gắn instrumentation vào các engine AI (như vLLM) bằng các OpenTelemetry span để ghi lại trạng thái bộ nhớ, độ dài sequence và chỉ số hiệu năng CUDA.

Cấu Hình OpenTelemetry Collector Daemon

Cấu hình OTel Collector để xử lý luồng telemetry throughput cao, trích xuất metadata (như batch size, tensor parallelism rank) thành metric label và lọc các nhiễu trace không cần thiết.

apiVersion: v1
kind: ConfigMap
metadata:
  name: otel-collector-config
  namespace: monitoring
data:
  otel-collector-config.yaml: |
    receivers:
      otlp:
        protocols:
          grpc:
            endpoint: 0.0.0.0:4317
          http:
            endpoint: 0.0.0.0:4318
      prometheus:
        config:
          scrape_configs:
            - job_name: 'nvidia-dcgm'
              scrape_interval: 2s
              static_configs:
                - targets: ['dcgm-exporter.gpu-operator.svc:9400']

    processors:
      memory_limiter:
        check_interval: 1s
        limit_percentage: 80
        spike_limit_percentage: 20
      batch:
        send_batch_size: 8192
        timeout: 1s
        send_batch_max_size: 10240
      transform:
        error_mode: ignore
        trace_statements:
          - context: span
            statements:
              - set(attributes["ai.vllm.gpu_cache_usage"], attributes["gpu.memory.used_ratio"])
              - set(attributes["ai.inference.ttft_ms"], attributes["http.response.time_first_token"])

    exporters:
      prometheus:
        endpoint: "0.0.0.0:8889"
        namespace: "otel_ai"
        send_timestamps: true
      otlp/jaeger:
        endpoint: "jaeger-collector.monitoring.svc:4317"
        tls:
          insecure: true

    service:
      pipelines:
        traces:
          receivers: [otlp]
          processors: [memory_limiter, transform, batch]
          exporters: [otlp/jaeger]
        metrics:
          receivers: [otlp, prometheus]
          processors: [memory_limiter, batch]
          exporters: [prometheus]

Bước 2: Xây Dựng Playbook Tự Động Hóa Với Robusta

Robusta đóng vai trò là control plane kết nối các sự kiện cảnh báo Kubernetes với các actor khắc phục. Dưới đây là mã Python custom action cho Robusta nhằm bắt cảnh báo AI infrastructure, trích xuất telemetry và gọi local LLM agent.

# robusta_ai_remediation.py
import requests
import json
import logging
from robusta.api import *

class AIIncidentPayload(BaseModel):
    alert_name: str
    pod_name: str
    namespace: str
    gpu_id: str
    error_logs: str
    otel_trace_id: str

@action
def ai_incident_llm_remediator(event: PrometheusKubernetesAlert, action_params: dict):
    """
    Robusta action được kích hoạt bởi sự kiện Alertmanager cho AI workloads.
    Trích xuất pod log, DCGM metric và OTel trace context, sau đó truy vấn local LLM agent.
    """
    pod = event.get_pod()
    if not pod:
        logging.error("Không tìm thấy Pod cho ngữ cảnh cảnh báo này.")
        return

    # Trích xuất log xung quanh thời điểm xảy ra lỗi
    pod_logs = event.get_pod_logs(num_lines=150)
    
    # Lấy OTel Trace ID từ label cảnh báo nếu có
    trace_id = event.alert.labels.get("otel_trace_id", "N/A")
    gpu_id = event.alert.labels.get("gpu_id", "0")

    payload = AIIncidentPayload(
        alert_name=event.alert_name,
        pod_name=pod.metadata.name,
        namespace=pod.metadata.namespace,
        gpu_id=gpu_id,
        error_logs=pod_logs,
        otel_trace_id=trace_id
    )

    # Gọi Local K8s LLM Agent
    llm_agent_url = action_params.get("agent_url", "http://ai-ops-agent.ai-system.svc:8080/v1/analyze")
    
    headers = {"Content-Type": "application/json"}
    try:
        response = requests.post(llm_agent_url, data=payload.json(), headers=headers, timeout=45)
        response.raise_for_status()
        execution_result = response.json()

        # Bổ sung kết quả phân tích & hành động của LLM vào Robusta Alert
        event.add_enrichment([
            MarkdownBlock(f"### 🤖 Phân Tích Sự Cố Tự Động Từ Local LLM
**Hành động đã thực hiện:** `{execution_result.get('action')}`

**Lý do:**
{execution_result.get('reasoning')}"),
            JsonBlock(json.dumps(execution_result.get("remediation_details")))
        ])
    except Exception as e:
        event.add_enrichment([MarkdownBlock(f"⚠️ **Kích hoạt AI Agent thất bại:** {str(e)}")])

Bước 3: Triển Khai Local LLM Agent Bảo Mật Nội Bộ

Local LLM Agent chạy trực tiếp trong cụm Kubernetes trên các GPU node chuyên dụng, xử lý ngữ cảnh vận hành an toàn mà không làm rò rỉ log hay trace ra ngoài. Controller sử dụng cơ chế function calling để tương tác chính xác với Kubernetes API.

# local_agent_server.py
import os
import requests
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel
from kubernetes import client, config

config.load_incluster_config()
k8s_v1 = client.CoreV1Api()
apps_v1 = client.AppsV1Api()

app = FastAPI(title="K8s AI-Ops Self-Healing Agent")

VLLM_ENDPOINT = os.getenv("VLLM_ENDPOINT", "http://vllm-service.ai-system.svc:8000/v1/chat/completions")

SYSTEM_PROMPT = """
Bạn là kĩ sư SRE chuyên gia về Hạ tầng AI Kubernetes và NVIDIA GPU Workloads.
Hãy phân tích payload sự cố (Alerts, DCGM metrics, OTel Spans, CUDA logs) và trả về JSON khắc phục sự cố.

CHỈ xuất ra JSON hợp lệ theo schema sau:
{
  "reasoning": "Phân tích chẩn đoán chi tiết",
  "action": "CHỌN_MỘT: [RESET_GPU, DYNAMIC_RESCALE_KV_CACHE, EVICT_POD, TAINT_NODE]",
  "remediation_details": {
    "target_resource": "string",
    "parameters": {}
  }
}
"""

class DiagnosticRequest(BaseModel):
    alert_name: str
    pod_name: str
    namespace: str
    gpu_id: str
    error_logs: str
    otel_trace_id: str

@app.post("/v1/analyze")
def analyze_and_remediate(req: DiagnosticRequest):
    user_prompt = f"""
    Cảnh báo sự cố: {req.alert_name}
    Pod mục tiêu: {req.namespace}/{req.pod_name} (GPU index: {req.gpu_id})
    OTel Trace ID: {req.otel_trace_id}
    
    Log gần nhất / Execution Trace:
    ```
    {req.error_logs}
    ```
    """

    vllm_payload = {
        "model": "deepseek-ai/DeepSeek-Coder-33B-instruct",
        "messages": [
            {"role": "system", "content": SYSTEM_PROMPT},
            {"role": "user", "content": user_prompt}
        ],
        "temperature": 0.0,
        "response_format": {"type": "json_object"}
    }

    res = requests.post(VLLM_ENDPOINT, json=vllm_payload, timeout=30)
    if res.status_code != 200:
        raise HTTPException(status_code=500, detail="Lỗi vLLM inference engine")

    llm_decision = res.json()["choices"][0]["message"]["content"]
    decision_json = json.loads(llm_decision)

    # Điều phối thực thi hành động
    action_status = execute_remediation_action(decision_json, req)
    decision_json["execution_status"] = action_status

    return decision_json

def execute_remediation_action(decision: dict, req: DiagnosticRequest) -> str:
    action = decision.get("action")
    
    if action == "EVICT_POD":
        k8s_v1.delete_namespaced_pod(name=req.pod_name, namespace=req.namespace)
        return f"Đã evicted thành công pod {req.pod_name}"
    
    elif action == "DYNAMIC_RESCALE_KV_CACHE":
        # Patch biến môi trường deployment để giảm tỉ lệ bộ nhớ KV cache
        dep = apps_v1.read_namespaced_deployment(name="vllm-inference", namespace=req.namespace)
        for container in dep.spec.template.spec.containers:
            if container.name == "vllm-container":
                for env in container.env:
                    if env.name == "GPU_MEMORY_UTILIZATION":
                        env.value = "0.80" # Giảm từ 0.90 xuống 0.80 để chống phân mảnh
        apps_v1.patch_namespaced_deployment(name="vllm-inference", namespace=req.namespace, body=dep)
        return "Đã patch vllm-inference GPU_MEMORY_UTILIZATION về 0.80"

    elif action == "TAINT_NODE":
        pod = k8s_v1.read_namespaced_pod(name=req.pod_name, namespace=req.namespace)
        node_name = pod.spec.node_name
        body = {
            "spec": {
                "taints": [{
                    "key": "gpu.nvidia.com/hardware-failure",
                    "value": "true",
                    "effect": "NoSchedule"
                }]
            }
        }
        k8s_v1.patch_node(node_name, body)
        return f"Đã gán taint NoSchedule cho node {node_name}"

    return "Không cần thực thi hành động."

Bước 4: Các Kịch Bản Khắc Phục Sự Cố Hạ Tầng AI Thực Tế

Bảng dưới đây tổng hợp các hành động tự động hóa tương ứng với từng lỗi hạ tầng GPU phổ biến:

Kịch Bản LỗiTín Hiệu OTel / DCGMĐiều Kiện Kích Hoạt Cảnh BáoHành Động Khắc Phục Của LLM Agent
CUDA KV-Cache OOMSpan phân bổ KV-cache vLLM quá tải (>98%) + `cudaErrorMemoryAllocation`Cảnh báo `vLLM_KVCacheExhausted` kéo dài 30sThực hiện patch rolling zero-downtime giảm tham số `gpu_memory_utilization`, đẩy bớt request quá tải sang node khác.
NCCL Synchronization TimeoutKhoảng trống trace trong span Ring-AllReduce > 60000msCảnh báo `NCCL_Communication_Timeout`Xác định GPU rank bị lỗi qua trace, evict pod worker hỏng và khởi tạo lại nhóm PyTorch Distributed Torchrun.
GPU ECC Uncorrectable Memory ErrorMetric DCGM `DCGM_FI_DEV_ECC_DBE_VOLATILE` > 0Cảnh báo `GPU_Hardware_ECC_Failure`Gán taint `NoSchedule` cho node hỏng, drain các pod đang chạy và kích hoạt daemonset reset driver GPU.
Vector Index Fragmented Latency SpikeĐộ trễ Vector Search Span > p99 2500ms + tranh chấp bộ nhớ caoCảnh báo `VectorDB_HighLatency_Tail`Thực thi job nén chỉ mục (index compaction) bất đồng bộ cho Qdrant/Milvus và dựng tạm pod read-replica.

Bước 5: Rào Cản An Toàn Và Quản Trị Doanh Nghiệp

Việc cho phép một AI Agent thực thi các lệnh quản trị lên cụm Kubernetes production đòi hỏi các cơ chế kiểm soát chặt chẽ:

  • Ép Kiểu JSON Schema Rõ Ràng: LLM bắt buộc phải xuất JSON hợp lệ thông qua các constrained sampling engine (như Guidance hoặc vLLM constrained decoding), từ chối các phản hồi dạng text tự do.
  • Rate-Limiting Cho Hành Động Khắc Phục: Robusta áp dụng cơ chế token-bucket (ví dụ: tối đa evict 3 pod trong 15 phút trên mỗi namespace) để tránh tình trạng restart dây chuyền toàn cụm.
  • Cơ Chế Duyệt Từ Con Người (Human-in-the-Loop): Đối với các hành động nguy hiểm (như drain node hoặc re-index storage), Robusta sẽ gửi kèm nút bấm phê duyệt trên Slack/Teams để SRE xác nhận trước khi thực hiện.
  • Ghi Log Kiểm Toán (Audit Logging): Mọi quyết định xử lý của LLM đều được lưu lại dưới dạng một OTel span riêng biệt, ghi nhận prompt, output của model và các sự kiện K8s API audit để phục vụ kiểm toán tuân thủ.

Kết Luận

Việc kết hợp OpenTelemetry tracing, Robusta event orchestration và khả năng suy luận của Local LLM giúp chuyển đổi hệ thống giám sát bị động truyền thống thành một hạ tầng AI tự chữa lành chủ động. Các đội ngũ Platform Engineering có thể đảm bảo tối đa độ khả dụng cho cụm GPU, loại bỏ các cuộc gọi xử lý sự cố ban đêm và duy trì tính bảo mật dữ liệu tuyệt đối nhờ vận hành toàn bộ trí tuệ nhân tạo bên trong ranh giới Kubernetes của doanh nghiệp.