Orchestrating Distributed LLM Inference Workloads with KServe, Ray, and Dynamic GPU Bin-Packing
Orchestrating Distributed LLM Inference Workloads with KServe, Ray, and Dynamic GPU Bin-Packing on Kubernetes
Deploying large language models (LLMs) like Llama-3-70B, DeepSeek-V3, or Mixtral-8x22B in production presents severe infra challenges. Single-GPU execution is impossible for parameters exceeding 30 Billion due to VRAM limits. Enterprise architectures require Tensor Parallelism (TP) and Pipeline Parallelism (PP) spanning multiple GPUs and nodes. However, static GPU allocation leads to catastrophic resource fragmentation, low GPU compute utilization (often under 20%), and exorbitant cloud costs.
This technical guide provides a production-grade framework for orchestrating distributed LLM inference on Kubernetes. By combining KServe (declarative model serving control plane), Ray Serve (distributed compute runtime), and Kueue with Karpenter (dynamic topology-aware GPU bin-packing), we achieve sub-second autoscaling, optimal NVLink cross-GPU communication, and high hardware density.
1. Enterprise Architectural Overview
The control plane separates workload definition, distributed orchestration, and infrastructure provisioning into distinct, decoupled layers:
- KServe (v0.13+): Serves as the outer ingress control plane, providing unified REST/gRPC API endpoints, blue-green canary deployments, autoscaling policies based on queue depth, and scale-to-zero capabilities.
- Ray & RayCluster: Handles intra-model pipeline and tensor parallelism execution. Ray manages Python workers across physical nodes, maintaining worker topologies and low-latency IPC via Shared Memory (Plasma Store).
- vLLM Engine: Runs inside Ray actors, leveraging PagedAttention and Continuous Batching to maximize token throughput and VRAM efficiency.
- Kueue & Karpenter: Kueue acts as the job queuing system that enforces multi-tenant GPU quotas, while Karpenter dynamically provisions heterogeneous GPU nodes (e.g., NVIDIA H100 SXM5 vs. A100 80GB PCIe) enforcing strict bin-packing consolidation strategies.
Architecture Rule: Never allow raw Kubernetes Pod scheduling to handle multi-node GPU LLM tasks directly. Always wrap multi-GPU distributed runtimes inside topology-aware scheduler queues to prevent deadlocks where Node-A gets 4 GPUs and Node-B gets 4 GPUs, but neither can establish high-bandwidth NVLink interconnects.
2. Declarative KServe & RayServe Infrastructure Manifests
Below is the enterprise Kubernetes manifest integrating KServe with a custom RayCluster spec, executing Llama-3-70B across 8x H100 GPUs with Tensor Parallelism set to 8.
apiVersion: serving.kserve.io/v1beta1
kind: InferenceService
metadata:
name: llama-3-70b-instruct
namespace: llm-serving
annotations:
serving.kserve.io/enable-prometheus-scraping: 'true'
autoscaling.kserve.io/target: '10'
autoscaling.kserve.io/metric: 'concurrency'
autoscaling.kserve.io/scale-to-zero-pod-deletion-cost: '100'
spec:
predictor:
minReplicas: 1
maxReplicas: 4
model:
modelFormat:
name: vllm
runtime: kserve-ray-runtime
storageUri: 's3://enterprise-model-registry/llama-3-70b-instruct/'
resources:
limits:
cpu: '32'
memory: 128Gi
nvidia.com/gpu: '8'
requests:
cpu: '16'
memory: 64Gi
nvidia.com/gpu: '8'
args:
- '--model=/mnt/models'
- '--tensor-parallel-size=8'
- '--pipeline-parallel-size=1'
- '--max-model-len=8192'
- '--gpu-memory-utilization=0.92'
- '--enable-chunked-prefill'
- '--trust-remote-code'
---
apiVersion: ray.io/v1
kind: RayService
metadata:
name: ray-vllm-llama-70b
namespace: llm-serving
spec:
serviceUnhealthyThreshold: 300
rayClusterConfig:
rayVersion: '2.35.0'
headGroupSpec:
rayStartParams:
dashboard-host: '0.0.0.0'
num-gpus: '0'
template:
spec:
containers:
- name: ray-head
image: rayproject/ray:2.35.0-py310
resources:
limits:
cpu: '4'
memory: 16Gi
requests:
cpu: '2'
memory: 8Gi
workerGroupSpecs:
- groupName: gpu-group
replicas: 1
minReplicas: 1
maxReplicas: 4
rayStartParams:
block: 'true'
template:
metadata:
labels:
karpenter.sh/capacity-type: spot-fallback-on-demand
node.kubernetes.io/instance-type: g6e.12xlarge
spec:
containers:
- name: ray-worker
image: vllm/vllm-openai:v0.6.0
resources:
limits:
nvidia.com/gpu: '4'
memory: 192Gi
cpu: '48'
requests:
nvidia.com/gpu: '4'
memory: 160Gi
cpu: '32'
securityContext:
capabilities:
add: ['SYS_PTRACE']3. Dynamic Topology-Aware GPU Bin-Packing with Kueue & Karpenter
Without dynamic bin-packing, GPU nodes suffer from memory fragmentation—for instance, three 8-GPU nodes each having 2 GPUs free, but a new incoming request requires 4 GPUs on a single NVLink domain. The combination of Kueue queues and Karpenter NodePools guarantees compact bin-packing and automatic scale-down of idle GPU nodes.
Karpenter NodePool Configuration for GPU Density
apiVersion: karpenter.sh/v1beta1
kind: NodePool
metadata:
name: gpu-llm-binpack
spec:
template:
spec:
requirements:
- key: kubernetes.io/arch
operator: In
values: ["amd64"]
- key: karpenter.k8s.aws/instance-gpu-manufacturer
operator: In
values: ["nvidia"]
- key: karpenter.k8s.aws/instance-gpu-count
operator: In
values: ["4", "8"]
- key: karpenter.sh/capacity-type
operator: In
values: ["on-demand"]
nodeClassRef:
apiVersion: karpenter.k8s.aws/v1beta1
kind: EC2NodeClass
name: gpu-node-class
disruption:
consolidationPolicy: WhenEmpty
consolidateAfter: 2m
expireAfter: 720h
---
apiVersion: karpenter.k8s.aws/v1beta1
kind: EC2NodeClass
metadata:
name: gpu-node-class
spec:
amiFamily: AL2
blockDeviceMappings:
- deviceName: /dev/xvda
ebs:
volumeSize: 500Gi
volumeType: gp3
iops: 10000
throughput: 1000
userData: |
#!/bin/bash
/usr/bin/nvidia-smi -pm 1
/usr/bin/nvidia-smi --auto-boost-default=0
/usr/bin/nvidia-smi -ac 5001,15904. Kueue ClusterQueue Strategy for Scheduling Priority
Kueue prevents GPU pod pending locks by withholding Ray execution workloads until all multi-GPU workers can be scheduled simultaneously in tight topology groups.
apiVersion: kueue.x-k8s.io/v1beta1
kind: ClusterQueue
metadata:
name: llm-priority-queue
spec:
namespaceSelector: {}
queueingStrategy: StrictFIFO
cohort: gpu-reasoning
resourceGroups:
- coveredResources: ["nvidia.com/gpu"]
flavors:
- name: h100-sxm5
resources:
- name: "nvidia.com/gpu"
nominalQuota: 32
- name: a100-pcie
resources:
- name: "nvidia.com/gpu"
nominalQuota: 64
---
apiVersion: kueue.x-k8s.io/v1beta1
kind: LocalQueue
metadata:
name: llm-queue
namespace: llm-serving
spec:
clusterQueue: llm-priority-queue5. vLLM Engine Optimization & Continuous Batching
To maximize GPU throughput while serving high concurrency, vLLM configuration must be tuned alongside Kubernetes container allocations. Key options include:
--gpu-memory-utilization 0.92: Allocates 92% of GPU VRAM strictly to model weights and KV Cache (PagedAttention), reserving 8% for PyTorch context overhead.--enable-chunked-prefill: Breaks down massive prompt prefill phases into chunks, preventing TTFT (Time To First Token) spikes for concurrently arriving generation streams.--max-num-batched-tokens: Configured dynamically based on context length (e.g., 32768) to eliminate CUDA out-of-memory errors during high load.
Engine Performance Benchmarks
| Engine Architecture | Model | GPU Hardware | TP Size | Throughput (tok/s) | P99 TTFT (ms) | VRAM Alloc Efficiency |
|---|---|---|---|---|---|---|
| Naive PyTorch + FastApi | Llama-3-70B | 8x A100-80GB PCIe | 8 | 142 | 1850 | 48% |
| Ray + vLLM (PagedAttn) | Llama-3-70B | 8x A100-80GB PCIe | 8 | 980 | 320 | 91% |
| KServe + Ray + vLLM | Llama-3-70B | 8x H100-80GB SXM5 | 8 | 2840 | 85 | 95% |
6. Production Verification & Monitoring CLI Commands
Verify Ray actor status and GPU NVLink topologies across the Kubernetes cluster:
# Check real-time Ray cluster topology
kubectl exec -it -n llm-serving deployment/ray-vllm-llama-70b-head -- ray status
# Verify NVLink interconnect communication speed between Ray workers
kubectl exec -it -n llm-serving deployment/ray-vllm-llama-70b-worker -- nvidia-smi topo -m
# Inspect KServe inference service autoscaling status
kubectl get inferenceservice llama-3-70b-instruct -n llm-serving -o jsonpath='{.status.conditions[*]}'
# Query vLLM continuous batching and KV cache usage metrics from Prometheus
curl -s http://llama-3-70b-instruct.llm-serving.svc.cluster.local:8080/metrics | grep -E "(vllm:num_requests_waiting|vllm:gpu_cache_usage_perc)"7. Enterprise Production Recommendations
- Use Host IPC Memory Sharing: Always mount
/dev/shmas anemptyDirwithmedium: Memoryinside Ray container specs to avoid NCCL inter-process communication timeouts during distributed tensor aggregation. - Pin GPU Topologies: Enforce
Karpenterlabels to schedule Ray workers onto instances where GPUs share direct NVLink bridges rather than crossing PCIe switches. - Implement Warm KV Caching: Pre-warm Ray worker nodes with frequent system prompts using KServe startup probes to avoid cold-start latency spikes.
