Back to articles
Technology Insight

Running DeepSeek-R1-Distill-Qwen-1.5B on a K3s Cluster: A Hybrid Edge AI Architecture with Raspberry Pi and ARM VPS

June 3, 2026

Introduction to Hybrid Edge AI Infrastructures

As large language models (LLMs) evolve, deploying them efficiently requires innovative infrastructure strategies. The rise of lightweight reasoning models, such as DeepSeek-R1-Distill-Qwen-1.5B, opens new possibilities for running advanced AI workloads outside traditional data centers. By combining local edge hardware with scalable cloud resources, businesses can achieve low latency, data privacy, and optimized cost efficiency.

This technical long-form post explores the architectural design and practical implementation of a hybrid Edge AI infrastructure. We will look at deploying a 1.5-billion-parameter reasoning model across a K3s Kubernetes cluster composed of three local Raspberry Pi single-board computers and one remote cloud-based ARM Virtual Private Server (VPS).

The Core Components: DeepSeek-R1, K3s, and ARM Architecture

To understand the strengths of this setup, it is essential to analyze the underlying technologies that make a hybrid edge-cloud deployment feasible for small-footprint AI models.

1. DeepSeek-R1-Distill-Qwen-1.5B

DeepSeek-R1 is a highly efficient reasoning model. The 1.5B variant distilled from Qwen2.5-Math retains sophisticated chain-of-thought capabilities while maintaining a compact physical footprint. Requiring fewer computing resources than its larger counterparts, this model can run on edge devices using quantization techniques (such as 4-bit or 8-bit GGUF formats) without a severe drop in mathematical or logical accuracy.

2. K3s (Lightweight Kubernetes)

Developed by SUSE/Rancher, K3s is a certified, lightweight Kubernetes distribution designed specifically for IoT, edge computing, and resource-constrained environments. Bundled into a single binary of less than 100MB, K3s eliminates unnecessary legacy drivers and alpha features, drastically reducing memory usage while providing standard Kubernetes API compliance.

3. The Hardware Stack

Our hybrid topology leverages two distinct compute tiers:

  • Edge Tier: Three Raspberry Pi single-board computers (ideally Raspberry Pi 4 or 5 with 8GB RAM). These nodes handle local preprocessing, secure local storage, and lightweight inference tasks.
  • Cloud Tier: One ARM64 VPS (such as an Ampere Altra-based instance on Oracle Cloud, AWS, or Akamai). This node acts as the resilient control plane and serves as a burst compute node for heavier processing tasks.

Architectural Blueprint: The Hybrid Network Topology

Deploying a single Kubernetes cluster across a mixture of a local area network (LAN) and a public cloud provider presents a primary challenge: secure and transparent node communication. Because edge devices usually sit behind residential or corporate NAT firewalls, they cannot naturally communicate with a cloud server over flat IP networks.

To solve this, K3s offers built-in support for establishing a secure mesh network across distributed nodes using WireGuard.

The diagram below conceptually demonstrates how control traffic and application queries flow through this decentralized system:

In this hybrid configuration, the cloud ARM VPS acts as the K3s Server (control plane) because it possesses a static public IP address and high availability. The three Raspberry Pi units run the K3s Agent (worker node) software, establishing outbound WireGuard connections to the cloud server. This creates a secure, encrypted overlay network where all pods can communicate directly, regardless of physical location.

Step-by-Step Deployment Strategy

Setting up this environment requires systematic configuration of the operating system, the networking layer, the container orchestration runtime, and the model execution engine.

Phase 1: Operating System Preparation

Before installing K3s, all nodes must be configured to support containerized workloads and appropriate memory allocation. On the Raspberry Pi nodes, you must enable legacy memory cgroups by editing the boot configuration file (/boot/firmware/cmdline.txt) and appending the following parameters:

cgroup_enable=cpuset cgroup_memory=1 cgroup_enable=memory

After saving the file, reboot the devices to apply the changes. This configuration prevents the K3s kubelet from throwing memory tracking errors during runtime.

Phase 2: Initializing the Cloud Control Plane

Execute the installation script on the ARM VPS to launch the master node. We explicitly pass the --flannel-backend=wireguard-native flag to instruct K3s to automatically create and manage the encrypted VPN mesh across the internet.

curl -sfL https://get.k3s.io | sh -s - server --flannel-backend=wireguard-native --node-external-ip=

Once initialized, extract the secure node registration token from the server. This token allows the edge agents to authenticate safely:

cat /var/lib/rancher/k3s/server/node-token

Phase 3: Connecting the Raspberry Pi Worker Nodes

On each of the three Raspberry Pi devices, run the agent installation script. Replace the environment variables with your specific VPS connection details:

curl -sfL https://get.k3s.io | K3S_URL=https://:6443 K3S_TOKEN= sh -

Verify the cluster health from your VPS or local machine via kubectl get nodes. You should see all four nodes listed in a Ready state, showing a mix of local edge and remote cloud architectures.

Optimizing DeepSeek-R1 Execution via Ollama or KubeRay

With the multi-node infrastructure running, the final step involves orchestrating the DeepSeek model. Because LLMs consume significant memory and CPU cycles on ARM processors, standard container execution requires optimization.

Two primary paradigms exist for serving the model within our K3s architecture:

  1. Ollama Daemon Deployment: Running Ollama as a Kubernetes Deployment or DaemonSet across the cluster. Ollama automatically manages GGUF quantizations, ensuring that the 1.5B model stays comfortably within the 4GB-8GB boundaries of individual Raspberry Pi nodes.
  2. KubeRay or vLLM Cluster: For advanced configurations, using a distributed framework like Ray allows you to split tensor parallelisms across nodes. However, due to network latency over the internet via WireGuard, running inference *locally inside single nodes* via localized Ollama deployments is highly recommended for hybrid setups to avoid cross-network bottlenecks.

Sample Kubernetes Manifest for Ollama Serving

Below is an optimized template for deploying the inference engine, target-scheduled specifically to the Raspberry Pi edge nodes using node affinity rules:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: deepseek-inference
spec:
  replicas: 3
  selector:
    matchLabels:
      app: deepseek-r1
  template:
    metadata:
      labels:
        app: deepseek-r1
    spec:
      affinity:
        nodeAffinity:
          requiredDuringSchedulingIgnoredDuringExecution:
            nodeSelectorTerms:
            - matchExpressions:
              - key: kubernetes.io/arch
                operator: In
                values:
                - arm64
      containers:
      - name: ollama
        image: ollama/ollama:latest
        resources:
          limits:
            memory: "4Gi"
            cpu: "3000m"
          requests:
            memory: "2Gi"
            cpu: "1500m"
        ports:
        - containerPort: 11434
        lifecycle:
          postStart:
            exec:
              command: ["/bin/sh", "-c", "ollama run deepseek-r1:1.5b"]

Performance Analytics and Business Benefits

Implementing a hybrid Edge AI infrastructure using small-scale components yields several distinct benefits for enterprises and engineering teams:

  • Cost Minimization: Utilizing inexpensive Raspberry Pi hardware reduces cloud infrastructure costs. The cloud VPS is used only for orchestration and high-level routing, dramatically dropping monthly API bills.
  • Data Privacy and Localization: Sensitive data can be processed on the local edge nodes without exposing information to public networks. Only sanitized or metadata queries need to hit the cloud layer.
  • Architectural Resilience: If the connection to the public internet drops, the local edge nodes can continue running local inference routines independently, guaranteeing high operational uptime for on-site automation.

Conclusion

The convergence of highly minimized reasoning architectures like DeepSeek-R1-Distill-Qwen-1.5B with resource-efficient orchestration platforms like K3s marks a shift in how we approach machine learning deployment. By designing a hybrid topology that links cloud flexibility with localized edge execution, organizations can deploy practical, private, and highly scalable AI tools today.

Running DeepSeek-R1-Distill-Qwen-1.5B on a K3s Cluster: A Hybrid Edge AI Architecture with Raspberry Pi and ARM VPS | DPTCloud