Back to articles
Technology Insight

Self-Hosting a Secure 'AI Code Interpreter' on Your VPS: Building a ChatGPT-Style Python Sandbox

May 25, 2026

Introduction: The Privacy Dilemma of AI Code Execution

The integration of Large Language Models (LLMs) with code execution environments—popularized by OpenAI's ChatGPT Advanced Data Analysis (formerly Code Interpreter)—has revolutionized data science, automation, and business intelligence. By allowing an LLM to write and execute Python code in real-time, organizations can automate complex data visualization, statistical analysis, and file manipulation. However, for enterprises handling proprietary data, financial records, or strictly regulated personal information, uploading data to third-party commercial platforms introduces significant compliance and security risks.

The solution lies in self-hosting. By deploying your own AI Code Interpreter infrastructure on a Virtual Private Server (VPS), you retain absolute control over your data lifecycle. This technical guide provides a comprehensive framework for architecting, securing, and deploying an isolated, containerized Python sandbox capable of executing LLM-generated code safely and efficiently.

The Architecture of a Secure Code Interpreter

To replicate the functionality of ChatGPT's Code Interpreter safely, your infrastructure must separate the brain (the LLM) from the muscle (the execution environment). Allowing an AI to run arbitrary code on a bare-metal server or a standard virtual machine invites severe security vulnerabilities, including remote code execution (RCE) breakouts, resource exhaustion, and unauthorized network access.

A robust self-hosted architecture relies on three primary layers:

  • The Orchestration Layer: Your application backend (e.g., built with LangChain, LlamaIndex, or custom Python code) that receives user prompts, queries the LLM, extracts the generated Python code blocks, and sends them to the sandbox.
  • The Gateway Layer: A secure API wrapper—typically utilizing Jupyter Kernel Gateway or headless Jupyter kernels—that receives code execution requests and passes them to an active Jupyter kernel.
  • The Isolation Layer (Sandbox): A highly restricted containerized environment (Docker or Podman) hardened to prevent host system contamination and resource hogging.

Step-by-Step Implementation Guide

Step 1: Setting Up the Hardened VPS Environment

Before deploying any containers, your host VPS must be hardened. Start by provisioning a clean Linux distribution (Ubuntu 24.04 LTS or Debian 12 are highly recommended) and configuring basic firewall rules using UFW (Uncomplicated Firewall) to block all unnecessary inbound traffic.

Critical Security Rule: Never run the Docker daemon or your orchestration application as the root user. Utilize Docker's rootless mode where possible, or ensure strict User Namespace remapping is enabled.

Step 2: Designing the Dockerized Python Sandbox

The core of the execution environment is a custom Docker image tailored for data science but stripped of dangerous privileges. Below is an optimized Dockerfile structure designed for this specific use case:

FROM python:3.11-slim

# Install essential data science dependencies
RUN pip install --no-cache-dir \
    jupyter_kernel_gateway \
    pandas \
    numpy \
    matplotlib \
    seaborn \
    scipy \
    scikit-learn

# Create a non-root user
RUN useradd -m -u 1001 sandboxuser
USER sandboxuser
WORKDIR /home/sandboxuser/workspace

# Expose Jupyter Kernel Gateway port
EXPOSE 8888

CMD ["jupyter", "kernelgateway", "--KernelGatewayApp.ip=0.0.0.0", "--KernelGatewayApp.port=8888", "--KernelGatewayApp.auth_token='YOUR_SECURE_TOKEN'"]

This configuration ensures that even if an LLM generates malicious code designed to wipe the file system, it is restricted to the non-root sandboxuser home directory inside an isolated container.

Step 3: Restricting Network Access and Resource Limits

An AI code interpreter should rarely require outbound internet access unless specifically instructed to scrape a website. To prevent data exfiltration (where malicious code transmits sensitive data to an external server), the sandbox container must be restricted using Docker network flags and strict resource constraints (cgroups).

When spinning up the container, execute it with defined resource ceilings to prevent Denial of Service (DoS) attacks caused by infinite loops or memory-intensive operations:

  1. Disable Networking: Use the --network none flag unless external API integration is mandatory. If network access is required, route traffic through an explicit internal bridge network with strict egress firewall rules.
  2. Limit CPU Utilization: Use --cpus="1.0" to prevent a single execution script from consuming 100% of the host VPS processing power.
  3. Limit Memory Allocation: Restrict the memory footprint via --memory="512m" to mitigate memory exhaustion vectors.
  4. Read-Only File System: Mount critical system directories as read-only, keeping only the specific /workspace directory writeable for data output and visualization rendering.

Connecting the LLM Orchestrator to the Sandbox

Once your Jupyter Kernel Gateway container is operational, your backend application needs to interact with it programmatically. When a user asks a data question, the process follows a strict execution loop:

First, the LLM processes the user query and generates a clean Python code block. Second, your application extracts this code block and submits it via a secure REST API or WebSocket connection to the Jupyter Kernel Gateway. Third, the gateway executes the code within the container, captures standard output (stdout), standard error (stderr), and any generated media assets (such as charts or graphs). Finally, the backend passes these results back to the LLM to format a final, human-readable response for the user.

Best Practices for Ongoing Production Security

Maintaining a secure self-hosted code interpreter requires vigilant monitoring and proactive infrastructure management. Implement the following practices to guarantee long-term stability:

Automated Container Recycling

Do not allow a single sandbox container to persist indefinitely. Implement a lifecycle manager that terminates, wipes, and recreates the Docker container after every user session or after a maximum idle timeout of 5 to 10 minutes. This guarantees that state corruption or residual temporary files do not carry over between distinct execution sessions.

Strict Execution Timeouts

Configure your orchestration backend to enforce a hard timeout (e.g., 30 seconds) on all code execution requests. If a script exceeds this window due to a complex calculation or an unoptimized loop, the execution thread should be forcefully terminated and an error returned to the model.

Comprehensive Audit Logging

Log every line of code generated by the LLM and executed within the sandbox. Store these logs on the host VPS system, entirely outside the reach of the container. In the event of an anomalous behavior pattern, these logs serve as an immutable audit trail for security reviews.

Conclusion

Deploying a self-hosted AI Code Interpreter on a VPS offers the ideal balance between modern AI capabilities and stringent data sovereignty. By isolating execution paths using Jupyter and Docker, enforcing strict resource ceilings, and cutting off untrusted network egress, you can confidently empower your business with advanced automated data analysis—without sacrificing security or privacy.

Self-Hosting a Secure 'AI Code Interpreter' on Your VPS: Building a ChatGPT-Style Python Sandbox | DPTCloud