Back to articles
Technology Insight

Orchestrating Distributed LLM Inference Workloads with KServe, Ray, and Dynamic GPU Bin-Packing

August 21, 2026

Orchestrating Distributed LLM Inference Workloads with KServe, Ray, and Dynamic GPU Bin-Packing on Kubernetes

Deploying large language models (LLMs) like Llama-3-70B, DeepSeek-V3, or Mixtral-8x22B in production presents severe infra challenges. Single-GPU execution is impossible for parameters exceeding 30 Billion due to VRAM limits. Enterprise architectures require Tensor Parallelism (TP) and Pipeline Parallelism (PP) spanning multiple GPUs and nodes. However, static GPU allocation leads to catastrophic resource fragmentation, low GPU compute utilization (often under 20%), and exorbitant cloud costs.

This technical guide provides a production-grade framework for orchestrating distributed LLM inference on Kubernetes. By combining KServe (declarative model serving control plane), Ray Serve (distributed compute runtime), and Kueue with Karpenter (dynamic topology-aware GPU bin-packing), we achieve sub-second autoscaling, optimal NVLink cross-GPU communication, and high hardware density.

1. Enterprise Architectural Overview

The control plane separates workload definition, distributed orchestration, and infrastructure provisioning into distinct, decoupled layers:

  • KServe (v0.13+): Serves as the outer ingress control plane, providing unified REST/gRPC API endpoints, blue-green canary deployments, autoscaling policies based on queue depth, and scale-to-zero capabilities.
  • Ray & RayCluster: Handles intra-model pipeline and tensor parallelism execution. Ray manages Python workers across physical nodes, maintaining worker topologies and low-latency IPC via Shared Memory (Plasma Store).
  • vLLM Engine: Runs inside Ray actors, leveraging PagedAttention and Continuous Batching to maximize token throughput and VRAM efficiency.
  • Kueue & Karpenter: Kueue acts as the job queuing system that enforces multi-tenant GPU quotas, while Karpenter dynamically provisions heterogeneous GPU nodes (e.g., NVIDIA H100 SXM5 vs. A100 80GB PCIe) enforcing strict bin-packing consolidation strategies.

Architecture Rule: Never allow raw Kubernetes Pod scheduling to handle multi-node GPU LLM tasks directly. Always wrap multi-GPU distributed runtimes inside topology-aware scheduler queues to prevent deadlocks where Node-A gets 4 GPUs and Node-B gets 4 GPUs, but neither can establish high-bandwidth NVLink interconnects.

2. Declarative KServe & RayServe Infrastructure Manifests

Below is the enterprise Kubernetes manifest integrating KServe with a custom RayCluster spec, executing Llama-3-70B across 8x H100 GPUs with Tensor Parallelism set to 8.

apiVersion: serving.kserve.io/v1beta1
kind: InferenceService
metadata:
  name: llama-3-70b-instruct
  namespace: llm-serving
  annotations:
    serving.kserve.io/enable-prometheus-scraping: 'true'
    autoscaling.kserve.io/target: '10'
    autoscaling.kserve.io/metric: 'concurrency'
    autoscaling.kserve.io/scale-to-zero-pod-deletion-cost: '100'
spec:
  predictor:
    minReplicas: 1
    maxReplicas: 4
    model:
      modelFormat:
        name: vllm
      runtime: kserve-ray-runtime
      storageUri: 's3://enterprise-model-registry/llama-3-70b-instruct/'
      resources:
        limits:
          cpu: '32'
          memory: 128Gi
          nvidia.com/gpu: '8'
        requests:
          cpu: '16'
          memory: 64Gi
          nvidia.com/gpu: '8'
      args:
        - '--model=/mnt/models'
        - '--tensor-parallel-size=8'
        - '--pipeline-parallel-size=1'
        - '--max-model-len=8192'
        - '--gpu-memory-utilization=0.92'
        - '--enable-chunked-prefill'
        - '--trust-remote-code'
---
apiVersion: ray.io/v1
kind: RayService
metadata:
  name: ray-vllm-llama-70b
  namespace: llm-serving
spec:
  serviceUnhealthyThreshold: 300
  rayClusterConfig:
    rayVersion: '2.35.0'
    headGroupSpec:
      rayStartParams:
        dashboard-host: '0.0.0.0'
        num-gpus: '0'
      template:
        spec:
          containers:
            - name: ray-head
              image: rayproject/ray:2.35.0-py310
              resources:
                limits:
                  cpu: '4'
                  memory: 16Gi
                requests:
                  cpu: '2'
                  memory: 8Gi
    workerGroupSpecs:
      - groupName: gpu-group
        replicas: 1
        minReplicas: 1
        maxReplicas: 4
        rayStartParams:
          block: 'true'
        template:
          metadata:
            labels:
              karpenter.sh/capacity-type: spot-fallback-on-demand
              node.kubernetes.io/instance-type: g6e.12xlarge
          spec:
            containers:
              - name: ray-worker
                image: vllm/vllm-openai:v0.6.0
                resources:
                  limits:
                    nvidia.com/gpu: '4'
                    memory: 192Gi
                    cpu: '48'
                  requests:
                    nvidia.com/gpu: '4'
                    memory: 160Gi
                    cpu: '32'
                securityContext:
                  capabilities:
                    add: ['SYS_PTRACE']

3. Dynamic Topology-Aware GPU Bin-Packing with Kueue & Karpenter

Without dynamic bin-packing, GPU nodes suffer from memory fragmentation—for instance, three 8-GPU nodes each having 2 GPUs free, but a new incoming request requires 4 GPUs on a single NVLink domain. The combination of Kueue queues and Karpenter NodePools guarantees compact bin-packing and automatic scale-down of idle GPU nodes.

Karpenter NodePool Configuration for GPU Density

apiVersion: karpenter.sh/v1beta1
kind: NodePool
metadata:
  name: gpu-llm-binpack
spec:
  template:
    spec:
      requirements:
        - key: kubernetes.io/arch
          operator: In
          values: ["amd64"]
        - key: karpenter.k8s.aws/instance-gpu-manufacturer
          operator: In
          values: ["nvidia"]
        - key: karpenter.k8s.aws/instance-gpu-count
          operator: In
          values: ["4", "8"]
        - key: karpenter.sh/capacity-type
          operator: In
          values: ["on-demand"]
      nodeClassRef:
        apiVersion: karpenter.k8s.aws/v1beta1
        kind: EC2NodeClass
        name: gpu-node-class
  disruption:
    consolidationPolicy: WhenEmpty
    consolidateAfter: 2m
    expireAfter: 720h
---
apiVersion: karpenter.k8s.aws/v1beta1
kind: EC2NodeClass
metadata:
  name: gpu-node-class
spec:
  amiFamily: AL2
  blockDeviceMappings:
    - deviceName: /dev/xvda
      ebs:
        volumeSize: 500Gi
        volumeType: gp3
        iops: 10000
        throughput: 1000
  userData: |
    #!/bin/bash
    /usr/bin/nvidia-smi -pm 1
    /usr/bin/nvidia-smi --auto-boost-default=0
    /usr/bin/nvidia-smi -ac 5001,1590

4. Kueue ClusterQueue Strategy for Scheduling Priority

Kueue prevents GPU pod pending locks by withholding Ray execution workloads until all multi-GPU workers can be scheduled simultaneously in tight topology groups.

apiVersion: kueue.x-k8s.io/v1beta1
kind: ClusterQueue
metadata:
  name: llm-priority-queue
spec:
  namespaceSelector: {}
  queueingStrategy: StrictFIFO
  cohort: gpu-reasoning
  resourceGroups:
    - coveredResources: ["nvidia.com/gpu"]
      flavors:
        - name: h100-sxm5
          resources:
            - name: "nvidia.com/gpu"
              nominalQuota: 32
        - name: a100-pcie
          resources:
            - name: "nvidia.com/gpu"
              nominalQuota: 64
---
apiVersion: kueue.x-k8s.io/v1beta1
kind: LocalQueue
metadata:
  name: llm-queue
  namespace: llm-serving
spec:
  clusterQueue: llm-priority-queue

5. vLLM Engine Optimization & Continuous Batching

To maximize GPU throughput while serving high concurrency, vLLM configuration must be tuned alongside Kubernetes container allocations. Key options include:

  • --gpu-memory-utilization 0.92: Allocates 92% of GPU VRAM strictly to model weights and KV Cache (PagedAttention), reserving 8% for PyTorch context overhead.
  • --enable-chunked-prefill: Breaks down massive prompt prefill phases into chunks, preventing TTFT (Time To First Token) spikes for concurrently arriving generation streams.
  • --max-num-batched-tokens: Configured dynamically based on context length (e.g., 32768) to eliminate CUDA out-of-memory errors during high load.

Engine Performance Benchmarks

Engine ArchitectureModelGPU HardwareTP SizeThroughput (tok/s)P99 TTFT (ms)VRAM Alloc Efficiency
Naive PyTorch + FastApiLlama-3-70B8x A100-80GB PCIe8142185048%
Ray + vLLM (PagedAttn)Llama-3-70B8x A100-80GB PCIe898032091%
KServe + Ray + vLLMLlama-3-70B8x H100-80GB SXM5828408595%

6. Production Verification & Monitoring CLI Commands

Verify Ray actor status and GPU NVLink topologies across the Kubernetes cluster:

# Check real-time Ray cluster topology
kubectl exec -it -n llm-serving deployment/ray-vllm-llama-70b-head -- ray status

# Verify NVLink interconnect communication speed between Ray workers
kubectl exec -it -n llm-serving deployment/ray-vllm-llama-70b-worker -- nvidia-smi topo -m

# Inspect KServe inference service autoscaling status
kubectl get inferenceservice llama-3-70b-instruct -n llm-serving -o jsonpath='{.status.conditions[*]}'

# Query vLLM continuous batching and KV cache usage metrics from Prometheus
curl -s http://llama-3-70b-instruct.llm-serving.svc.cluster.local:8080/metrics | grep -E "(vllm:num_requests_waiting|vllm:gpu_cache_usage_perc)"

7. Enterprise Production Recommendations

  1. Use Host IPC Memory Sharing: Always mount /dev/shm as an emptyDir with medium: Memory inside Ray container specs to avoid NCCL inter-process communication timeouts during distributed tensor aggregation.
  2. Pin GPU Topologies: Enforce Karpenter labels to schedule Ray workers onto instances where GPUs share direct NVLink bridges rather than crossing PCIe switches.
  3. Implement Warm KV Caching: Pre-warm Ray worker nodes with frequent system prompts using KServe startup probes to avoid cold-start latency spikes.