Multi-Cloud Cross-Cluster Networking with Cilium ClusterMesh, BGP, and eBPF
Architecting Multi-Cloud Cross-Cluster Networking with Cilium ClusterMesh, BGP, and eBPF
As enterprise Kubernetes environments expand across multi-cloud infrastructure (e.g., AWS EKS, Google Cloud GKE, and Bare-Metal/On-Premises), traditional networking models fail to meet stringency demands around low latency, high throughput, predictable IP routing, and strict zero-trust security. Traditional overlay networks (such as VXLAN or Geneve without native routing) introduce packet encapsulation overhead, restrict MTU sizes, increase CPU usage, and obscure pod IP visibility from external network appliances.
Combining Cilium ClusterMesh, eBPF-driven Native Routing, and the Cilium BGP Control Plane solves these bottlenecks. This architecture enables seamless, high-performance, pod-to-pod cross-cluster routing across cloud boundaries without requiring sidecar proxies or double-NATing gateway appliances.
1. Enterprise Architectural Overview
The solution relies on three foundational components operating across separate cloud environments:
- eBPF Datapath: Replaces standard
iptablesand Linux bridge networking with kernel-level eBPF programs attached to TC (Traffic Control) ingress/egress filters and XDP hooks. This provides socket-level load balancing, direct packet routing, and Transparent Encryption via WireGuard or IPsec. - Cilium BGP Control Plane: Dynamically advertises Kubernetes Pod CIDRs, Service LoadBalancer IPs, and ClusterIPs to top-of-rack (ToR) switches or cloud provider virtual routers (AWS Transit Gateway, GCP Cloud Router) using standard BGP peering protocols (GoBGP integrated into the Cilium agent).
- Cilium ClusterMesh: Synchronizes Kubernetes service endpoints and identity metadata across multiple independent clusters using an external or embedded etcd KVStore mesh. This establishes global multi-cluster service discovery, automatic failover, and unified network security policies.
+---------------------------------------+ +---------------------------------------+
| AWS EKS Cluster | | GCP GKE Cluster |
| Pod CIDR: 10.240.0.0/16 (Cluster ID:1) | | Pod CIDR: 10.241.0.0/16 (Cluster ID:2)|
| +---------------------------------+ | | +---------------------------------+ |
| | Cilium Agent + eBPF | | | | Cilium Agent + eBPF | |
| +---------------------------------+ | | +---------------------------------+ |
| | BGP Session (eBGP) | | | BGP Session (eBGP) |
+---------|-----------------------------+ +---------|-----------------------------+
| |
v v
+------------------+ AWS DirectConnect / Cloud Interconnect +------------------+
| AWS Transit Gateway| <==========================================> | GCP Cloud Router |
+------------------+ +------------------+2. Prerequisites & IPAM Strategy
Cross-cluster native routing with BGP requires a non-overlapping IP address management (IPAM) blueprint. Pod CIDRs and Node CIDRs across all cloud providers must be uniquely allocatable.
| Cluster Context | Cluster ID | Node CIDR | Pod CIDR | Service CIDR |
|---|---|---|---|---|
| AWS EKS (us-east-1) | 1 | 10.100.0.0/16 | 10.240.0.0/16 | 172.20.0.0/16 |
| GCP GKE (us-central1) | 2 | 10.101.0.0/16 | 10.241.0.0/16 | 172.21.0.0/16 |
3. Configuring Cilium with Native Routing and BGP
Deploy Cilium using Helm, explicitly enabling native routing mode (disabling encapsulation like VXLAN), allocating unique Cluster IDs, and activating the BGP control plane.
# cilium-aws-values.yaml
cluster:
name: eks-us-east-1
id: 1
routingMode: native
ipv4NativeRoutingCIDR: 10.0.0.0/8
autoDirectNodeRoutes: true
ebpf:
masquerade: true
bgpControlPlane:
enabled: true
clustermesh:
useAPIServer: true
apiserver:
service:
type: LoadBalancer
ipam:
mode: kubernetes
tunnelMode: "disabled"
endpointRoutes:
enabled: trueAfter deploying Cilium via Helm, apply the CiliumBGPPeeringPolicy to establish eBGP peering with the local cloud router (e.g., AWS Transit Gateway Appliance or Quagga/BIRD virtual router).
apiVersion: cilium.io/v2alpha1
kind: CiliumBGPPeeringPolicy
metadata:
name: bgp-peering-aws
spec:
nodeSelector:
matchLabels:
cilium.io/bgp-virtual-router: "true"
virtualRouters:
- localASN: 65001
exportPodCIDR: true
neighbors:
- peerAddress: "10.100.0.1/32"
peerASN: 65000
eBGPMultihop: 1
connectRetryTimeSeconds: 10
holdTimeSeconds: 30
keepAliveTimeSeconds: 10
serviceSelector:
matchLabels:
type: global-lb4. Establishing Cilium ClusterMesh across Clouds
Once native IP connectivity is verified via BGP routes between AWS and GCP networks, initialize ClusterMesh across both clusters using the Cilium CLI.
# Enable ClusterMesh on AWS Cluster
cilium clustermesh enable --context context-aws-eks --service-type LoadBalancer
# Enable ClusterMesh on GCP Cluster
cilium clustermesh enable --context context-gcp-gke --service-type LoadBalancer
# Connect AWS to GCP ClusterMesh
cilium clustermesh connect --context context-aws-eks --destination-context context-gcp-gke
# Validate Connection Status
cilium clustermesh status --context context-aws-eks5. Global Service Discovery & Policy Enforcement
To enable cross-cluster load balancing, define a Kubernetes Service with the service.cilium.io/global: "true" annotation. EndpointSlices are automatically shared across clusters via the ClusterMesh KVStore.
apiVersion: v1
kind: Service
metadata:
name: payment-gateway
namespace: production
annotations:
service.cilium.io/global: "true"
service.cilium.io/shared: "true"
service.cilium.io/global-affinity: "local"
spec:
type: ClusterIP
ports:
- port: 8080
targetPort: 8080
protocol: TCP
selector:
app: payment-gatewayWith eBPF datapath integration, cross-cluster security policies are enforced seamlessly using Kubernetes label selectors and Cilium security identities, regardless of underlying cloud boundaries:
apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
name: allow-cross-cluster-checkout
namespace: production
spec:
endpointSelector:
matchLabels:
app: payment-gateway
ingress:
- fromEndpoints:
- matchLabels:
app: checkout-service
io.cilium.k8s.policy.cluster: gke-us-central1
toPorts:
- ports:
- port: "8080"
protocol: TCP6. Verification and Low-Level Diagnostics
Validate eBPF map state and BGP route propagation using native Cilium debugging tools directly inside the agent pods:
# Verify BGP peering state
kubectl -n kube-system exec -it ds/cilium -- cilium-dbg bgp peers
# Inspect eBPF ipcache for cross-cluster pod IP to identity mapping
kubectl -n kube-system exec -it ds/cilium -- cilium-dbg map get cilium_ipcache
# Trace cross-cluster packet traversal via eBPF monitor
kubectl -n kube-system exec -it ds/cilium -- monitor --type drop,bpf7. Conclusion
Architecting multi-cloud cross-cluster networking with Cilium ClusterMesh, BGP, and eBPF eliminates performance bottlenecks associated with legacy overlay VPNs and sidecar proxies. By driving routing decisions directly in the Linux kernel via eBPF and advertising non-overlapping CIDRs via standard BGP, enterprises achieve bare-metal packet processing speeds, robust zero-trust security, and high-availability global service resilience across diverse cloud environments.
