Architecting Carbon-Aware Kubernetes Workloads with Kepler and KEDA
Architecting Carbon-Aware Kubernetes Workloads with Kepler and KEDA
Introduction
As enterprise cloud footprints expand, sustainable software engineering—commonly referred to as GreenOps—has evolved from a corporate social responsibility initiative into a core architectural requirement. Traditional Kubernetes autoscalers scale workloads strictly based on infrastructure performance indicators such as CPU and memory utilization. However, this approach ignores the environmental cost: the carbon intensity of the underlying electricity grid fluctuates dynamically based on weather conditions, time of day, and regional grid mixes.
To build truly sustainable cloud-native systems, software architects must shift toward carbon-aware computing. By combining Kepler (Kubernetes-based Efficient Power Level Exporter) with KEDA (Kubernetes Event-driven Autoscaling), organizations can measure actual container-level energy consumption and dynamically scale down non-critical workloads during periods of high grid carbon intensity, or shift workloads to regions utilizing cleaner energy.
Core Value Proposition
Transitioning to a carbon-aware architecture yields distinct operational and environmental benefits:
-
Granular Energy Telemetry: Leverage eBPF to map power consumption directly to individual Pods, Namespaces, and Containers without code instrumentation.
-
Grid-Responsive Scaling: Automatically defer batch computing, machine learning training, or heavy analytical processing to periods of high renewable energy availability.
-
Optimized Cloud Spend: GreenOps aligns naturally with FinOps; running workloads when energy is cleaner often correlates with lower off-peak pricing on spot instances.
-
Regulatory Compliance: Generate precise Scope 3 carbon emission reports for auditability under emerging global environmental reporting standards.
Architecture & System Design
The carbon-aware autoscaling loop is comprised of three key architectural layers:
+-------------------------------------------------------------+ | Carbon API | | (Electricity Maps / Carbon Intensity) | +------------------------------+------------------------------+ | v +-------------------------------------------------------------+ | Prometheus / Grafana | | (Aggregates Kepler Energy & Carbon Metrics) | +------------------------------+------------------------------+ ^ | (Queries Metrics) +------------------------------+------------------------------+ | KEDA | | (Manages Pod Scaling via Custom ScaledObject) | +------------------------------+------------------------------+ | v +-------------------------------------------------------------+ | Kubernetes Pod / HPA | +-------------------------------------------------------------+ ^ ^ | (Profiles Power via eBPF) | (Monitors) +-------------------------------+ +------+-------+ | Kepler DaemonSet | ------------>| Node Power | +-------------------------------+ +--------------+
1. The Telemetry Layer (Kepler)
Kepler runs as a DaemonSet across the cluster. It hooks into the Linux kernel via eBPF to trace CPU instructions, hardware performance counters, and cache misses. Kepler then feeds these performance indicators into a regression model alongside system-level energy metrics (retrieved from Intel RAPL or ACPI) to accurately calculate the energy footprint of each container in real time.
2. The Aggregation Layer (Prometheus)
Prometheus scrapes the power metrics exposed by Kepler (e.g., kepler_container_joules_total). Simultaneously, a custom exporter or Prometheus collector queries external Carbon Intensity APIs (such as WattTime or Electricity Maps) to pull the current regional grid carbon intensity in grams of CO2 equivalent per kilowatt-hour ($gCO_2eq/kWh$).
3. The Orchestration Layer (KEDA)
KEDA continuously monitors the Prometheus server. By defining a custom ScaledObject that targets carbon-aware PromQL metrics, KEDA coordinates scaling decisions. When carbon intensity exceeds a configured threshold, KEDA downscales non-essential workloads to their minimum replica limits, ramping them back up once the grid relies on cleaner energy sources.
Technical Deep Dive & Deployment
Step 1: Deploying Kepler
First, deploy Kepler to your cluster using Helm. Kepler requires access to kernel headers to compile and load its eBPF programs.
helm repo add kepler-exporter https://sustainable-computing-io.github.io/kepler-helm-chart
helm repo update
helm install kepler kepler-exporter/kepler \
--namespace kepler --create-namespace \
--set serviceMonitor.enabled=true
Step 2: Designing the Carbon-Aware Scaling Logic
To demonstrate carbon-aware autoscaling, let's configure a batch-processing system. We want this workload to run at high capacity (e.g., 10 replicas) when the carbon intensity is low (under $150\ gCO_2eq/kWh$), but scale down to 1 replica when the grid is dirty.
We assume Prometheus exposes a metric called grid_carbon_intensity_gco2_per_kwh. We can construct a PromQL query that acts as the scaling trigger:
promql sum(grid_carbon_intensity_gco2_per_kwh{region="us-east"})
Step 3: Configuring the KEDA ScaledObject
Apply the following KEDA ScaledObject manifest. It targets a deployment named batch-processor and scales its replicas based on carbon intensity thresholds:
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: carbon-aware-batch-scaler
namespace: processing
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: batch-processor
minReplicaCount: 1
maxReplicaCount: 10
cooldownPeriod: 300
advanced:
horizontalPodAutoscalerConfig:
behavior:
scaleDown:
stabilizationWindowSeconds: 600
policies:
- type: Percent
value: 20
periodSeconds: 60
triggers:
- type: prometheus metadata: serverAddress: http://prometheus-k8s.monitoring.svc.cluster.local:9090 metricName: grid_carbon_intensity_gco2_per_kwh # We scale DOWN when carbon intensity is HIGH. # Mathematically, we set a threshold where higher values target fewer pods. query: | vector(10) - (clamp_min(grid_carbon_intensity_gco2_per_kwh{region="us-east"} - 150, 0) / 50) threshold: '1'
Architectural Note: The custom PromQL calculation above is designed to degrade gracefully. As grid carbon intensity climbs past the $150\ gCO_2eq/kWh$ baseline, the targeted metric value drops, triggering KEDA to gradually reduce the replica target of the deployment.
Enterprise Security & Best Practices
-
Least Privilege eBPF: Kepler requires privileged security contexts to load eBPF programs into the kernel. Restrict Kepler's deployment strictly to designated system namespaces using Kubernetes Pod Security Standards (PSS) and ensure DaemonSets are monitored for unauthorized modifications.
-
Resiliency and Fail-Safes: Carbon APIs can experience downtime or rate limiting. Ensure your PromQL scaling queries fallback to a default safe value using the
keep_lastordefaultoperators, preventing workloads from scaling down to zero during API outages. -
Coordinated Descheduling: Use KEDA's
cooldownPeriodand HPA stabilization windows to prevent scaling thrashing when the grid carbon intensity oscillates rapidly near the threshold bounds.
Conclusion
By integrating eBPF-powered energy profiling through Kepler with event-driven autoscaling via KEDA, enterprises can build highly responsive, carbon-aware infrastructures. This architecture not only reduces a company's environmental impact but also establishes a sustainable foundation for running high-throughput, cloud-native systems. Designing for the greenest path ensures resource optimization, reduced cloud expenditures, and compliance with tomorrow's strict environmental regulations.
