Running DeepSeek-R1-Distill-Qwen-1.5B on a K3s Cluster: A Hybrid Edge AI Architecture with Raspberry Pi and ARM VPS
Introduction to Hybrid Edge AI Infrastructures
As large language models (LLMs) evolve, deploying them efficiently requires innovative infrastructure strategies. The rise of lightweight reasoning models, such as DeepSeek-R1-Distill-Qwen-1.5B, opens new possibilities for running advanced AI workloads outside traditional data centers. By combining local edge hardware with scalable cloud resources, businesses can achieve low latency, data privacy, and optimized cost efficiency.
This technical long-form post explores the architectural design and practical implementation of a hybrid Edge AI infrastructure. We will look at deploying a 1.5-billion-parameter reasoning model across a K3s Kubernetes cluster composed of three local Raspberry Pi single-board computers and one remote cloud-based ARM Virtual Private Server (VPS).
The Core Components: DeepSeek-R1, K3s, and ARM Architecture
To understand the strengths of this setup, it is essential to analyze the underlying technologies that make a hybrid edge-cloud deployment feasible for small-footprint AI models.
1. DeepSeek-R1-Distill-Qwen-1.5B
DeepSeek-R1 is a highly efficient reasoning model. The 1.5B variant distilled from Qwen2.5-Math retains sophisticated chain-of-thought capabilities while maintaining a compact physical footprint. Requiring fewer computing resources than its larger counterparts, this model can run on edge devices using quantization techniques (such as 4-bit or 8-bit GGUF formats) without a severe drop in mathematical or logical accuracy.
2. K3s (Lightweight Kubernetes)
Developed by SUSE/Rancher, K3s is a certified, lightweight Kubernetes distribution designed specifically for IoT, edge computing, and resource-constrained environments. Bundled into a single binary of less than 100MB, K3s eliminates unnecessary legacy drivers and alpha features, drastically reducing memory usage while providing standard Kubernetes API compliance.
3. The Hardware Stack
Our hybrid topology leverages two distinct compute tiers:
- Edge Tier: Three Raspberry Pi single-board computers (ideally Raspberry Pi 4 or 5 with 8GB RAM). These nodes handle local preprocessing, secure local storage, and lightweight inference tasks.
- Cloud Tier: One ARM64 VPS (such as an Ampere Altra-based instance on Oracle Cloud, AWS, or Akamai). This node acts as the resilient control plane and serves as a burst compute node for heavier processing tasks.
Architectural Blueprint: The Hybrid Network Topology
Deploying a single Kubernetes cluster across a mixture of a local area network (LAN) and a public cloud provider presents a primary challenge: secure and transparent node communication. Because edge devices usually sit behind residential or corporate NAT firewalls, they cannot naturally communicate with a cloud server over flat IP networks.
To solve this, K3s offers built-in support for establishing a secure mesh network across distributed nodes using WireGuard.
The diagram below conceptually demonstrates how control traffic and application queries flow through this decentralized system:
In this hybrid configuration, the cloud ARM VPS acts as the K3s Server (control plane) because it possesses a static public IP address and high availability. The three Raspberry Pi units run the K3s Agent (worker node) software, establishing outbound WireGuard connections to the cloud server. This creates a secure, encrypted overlay network where all pods can communicate directly, regardless of physical location.
Step-by-Step Deployment Strategy
Setting up this environment requires systematic configuration of the operating system, the networking layer, the container orchestration runtime, and the model execution engine.
Phase 1: Operating System Preparation
Before installing K3s, all nodes must be configured to support containerized workloads and appropriate memory allocation. On the Raspberry Pi nodes, you must enable legacy memory cgroups by editing the boot configuration file (/boot/firmware/cmdline.txt) and appending the following parameters:
cgroup_enable=cpuset cgroup_memory=1 cgroup_enable=memory
After saving the file, reboot the devices to apply the changes. This configuration prevents the K3s kubelet from throwing memory tracking errors during runtime.
Phase 2: Initializing the Cloud Control Plane
Execute the installation script on the ARM VPS to launch the master node. We explicitly pass the --flannel-backend=wireguard-native flag to instruct K3s to automatically create and manage the encrypted VPN mesh across the internet.
curl -sfL https://get.k3s.io | sh -s - server --flannel-backend=wireguard-native --node-external-ip=
Once initialized, extract the secure node registration token from the server. This token allows the edge agents to authenticate safely:
cat /var/lib/rancher/k3s/server/node-token
Phase 3: Connecting the Raspberry Pi Worker Nodes
On each of the three Raspberry Pi devices, run the agent installation script. Replace the environment variables with your specific VPS connection details:
curl -sfL https://get.k3s.io | K3S_URL=https://
Verify the cluster health from your VPS or local machine via kubectl get nodes. You should see all four nodes listed in a Ready state, showing a mix of local edge and remote cloud architectures.
Optimizing DeepSeek-R1 Execution via Ollama or KubeRay
With the multi-node infrastructure running, the final step involves orchestrating the DeepSeek model. Because LLMs consume significant memory and CPU cycles on ARM processors, standard container execution requires optimization.
Two primary paradigms exist for serving the model within our K3s architecture:
- Ollama Daemon Deployment: Running Ollama as a Kubernetes Deployment or DaemonSet across the cluster. Ollama automatically manages GGUF quantizations, ensuring that the 1.5B model stays comfortably within the 4GB-8GB boundaries of individual Raspberry Pi nodes.
- KubeRay or vLLM Cluster: For advanced configurations, using a distributed framework like Ray allows you to split tensor parallelisms across nodes. However, due to network latency over the internet via WireGuard, running inference *locally inside single nodes* via localized Ollama deployments is highly recommended for hybrid setups to avoid cross-network bottlenecks.
Sample Kubernetes Manifest for Ollama Serving
Below is an optimized template for deploying the inference engine, target-scheduled specifically to the Raspberry Pi edge nodes using node affinity rules:
apiVersion: apps/v1
kind: Deployment
metadata:
name: deepseek-inference
spec:
replicas: 3
selector:
matchLabels:
app: deepseek-r1
template:
metadata:
labels:
app: deepseek-r1
spec:
affinity:
nodeAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms:
- matchExpressions:
- key: kubernetes.io/arch
operator: In
values:
- arm64
containers:
- name: ollama
image: ollama/ollama:latest
resources:
limits:
memory: "4Gi"
cpu: "3000m"
requests:
memory: "2Gi"
cpu: "1500m"
ports:
- containerPort: 11434
lifecycle:
postStart:
exec:
command: ["/bin/sh", "-c", "ollama run deepseek-r1:1.5b"]Performance Analytics and Business Benefits
Implementing a hybrid Edge AI infrastructure using small-scale components yields several distinct benefits for enterprises and engineering teams:
- Cost Minimization: Utilizing inexpensive Raspberry Pi hardware reduces cloud infrastructure costs. The cloud VPS is used only for orchestration and high-level routing, dramatically dropping monthly API bills.
- Data Privacy and Localization: Sensitive data can be processed on the local edge nodes without exposing information to public networks. Only sanitized or metadata queries need to hit the cloud layer.
- Architectural Resilience: If the connection to the public internet drops, the local edge nodes can continue running local inference routines independently, guaranteeing high operational uptime for on-site automation.
Conclusion
The convergence of highly minimized reasoning architectures like DeepSeek-R1-Distill-Qwen-1.5B with resource-efficient orchestration platforms like K3s marks a shift in how we approach machine learning deployment. By designing a hybrid topology that links cloud flexibility with localized edge execution, organizations can deploy practical, private, and highly scalable AI tools today.
