Back to articles
Technology Insight

Implementing Open-WebUI Pipelines: Building Custom Guardrail Filters to Prevent Malware and Data Leaks in Enterprise AI

June 2, 2026

Introduction: The Enterprise AI Security Dilemma

As enterprises rapidly adopt Large Language Models (LLMs) to drive productivity, security teams face a dual challenge: intellectual property data leaks and the injection of malicious code. While open-source interfaces like Open-WebUI offer incredible flexibility for corporate deployments, standard installations lack the granular, business-specific compliance controls required by enterprise risk management frameworks.

To bridge this gap, Open-WebUI introduced Pipelines—a plugin architecture that intercepts inputs and outputs before they reach the user or the underlying LLM. In this comprehensive guide, we will explore how to design, write, and deploy a custom Python-based Guardrails filter within Open-WebUI Pipelines to scan for malware strings and prevent sensitive data exfiltration.

Understanding Open-WebUI Pipelines Architecture

Before diving into the code, it is essential to understand where Pipelines sit within your AI infrastructure. Open-WebUI Pipelines operate as an independent middleware layer. When a user submits a prompt, it does not go directly to the LLM; instead, it passes through the Pipeline framework.

Key Architecture Flow: User Prompt → Pipeline Filter (Inlet) → LLM Inference → Pipeline Filter (Outlet) → User Interface.

By leveraging this architecture, developers can implement two critical security checkpoints:

  • Inlet Filters: Scans the user's prompt for malicious intent, unauthorized code, or Personally Identifiable Information (PII) before it hits the LLM API.
  • Outlet Filters: Evaluates the LLM's response to ensure it does not accidentally reveal proprietary code, internal API keys, or restricted business metrics.

Setting Up Your Pipeline Environment

To implement custom filters, you must first ensure that your Open-WebUI instance is connected to the Pipelines gateway. Pipelines usually run in a separate Docker container to maintain isolation and prevent rogue scripts from compromising the main web interface.

Step 1: Deploying the Pipelines Container

Run the following Docker command to spin up the official Pipelines environment alongside your existing setup:

docker run -d -p 9099:9099 --add-host=host.docker.internal:host-gateway -v pipelines:/app/pipelines --name pipelines ghcr.io/open-webui/pipelines:main

Step 2: Connecting Open-WebUI to Pipelines

  1. Navigate to the Admin Settings panel in your Open-WebUI interface.
  2. Select the Connections or Pipelines section.
  3. Enter the gateway URL: http://localhost:9099 (or the internal network IP of your container).

Designing the Custom Guardrail Filter

Our enterprise guardrail will focus on two major compliance vectors: Data Leakage Prevention (DLP) and Malicious Code Detection. We will write a Python script utilizing regular expressions (Regex) and basic text classification signatures to block high-risk actions.

Core Components of a Valve Class

Open-WebUI utilizes a configuration concept called Valves. Valves allow administrators to adjust settings (like blocked keywords or regex rules) directly from the UI without modifying the underlying Python code. This provides non-technical compliance officers the ability to update security rules on the fly.

The Source Code: enterprise_guardrail.py

Below is the complete, production-ready Python implementation for the enterprise guardrail pipeline. This script must be placed in the /app/pipelines directory of your container.

from typing import List, Optional, Union, Dict
import re
from pydantic import BaseModel, Field

class Pipeline:
    class Valves(BaseModel):
        # Configuration settings exposed to the Open-WebUI Admin UI
        block_patterns: List[str] = Field(
            default=[
                r"(?i)CONFIDENTIAL_DO_NOT_DISTRIBUTE",
                r"(4[0-9]{12}(?:[0-9]{3})?)", # Generic Visa Credit Card
                r"AI_SECRET_[a-zA-Z0-9]{32}" # Mock Corporate API Key
            ],
            description="Regex patterns to detect data leaks and corporate secrets."
        )
        malware_signatures: List[str] = Field(
            default=[
                "eval(base64_decode",
                "powershell.exe -nop -w hidden -c",
                "/etc/passwd"
            ],
            description="Common code injection and malware strings."
        )
        rejection_message: str = Field(
            default="Security Alert: Your request violates corporate data protection policy.",
            description="Message returned to the user when a block occurs."
        )

    def __init__(self):
        self.name = "Enterprise Guardrails Filter"
        self.valves = self.Valves()

    async def inlet(self, body: dict, user: Optional[dict] = None) -> dict:
        # Access the user's latest prompt
        messages = body.get("messages", [])
        if not messages:
            return body

        last_message_content = messages[-1].get("content", "")

        # 1. Check for Data Leaks via Regex Valves
        for pattern in self.valves.block_patterns:
            if re.search(pattern, last_message_content):
                raise Exception(f"{self.valves.rejection_message} (Triggered by Policy: Data Protection)")

        # 2. Check for Malicious Code/Strings
        for signature in self.valves.malware_signatures:
            if signature in last_message_content:
                raise Exception(f"{self.valves.rejection_message} (Triggered by Policy: Malware Prevention)")

        return body

    async def outlet(self, body: dict, user: Optional[dict] = None) -> dict:
        # Optional: Apply similar compliance rules to the LLM\'s output
        return body

Testing and Validating the Guardrails

Once you have saved the file into the pipelines directory, restart the pipeline service or trigger a reload from the admin interface. It is time to validate that our policy rules are functioning properly.

Test Case 1: Data Leakage Prevention

Attempt to paste a mock API key into the chat window:

User prompt: "Can you optimize this function? The token is AI_SECRET_a1b2c3d4e5f6g7h8i9j0k1l2m3n4o5p6"

Expected Outcome: The UI immediately intercepts the request and displays: Security Alert: Your request violates corporate data protection policy. (Triggered by Policy: Data Protection). The prompt never reaches OpenAI, Anthropic, or your local Llama instance.

Test Case 2: Malware String Filtering

Attempt to ask the model to analyze a malicious PowerShell script:

User prompt: "What does this script do? powershell.exe -nop -w hidden -c IEX (New-Object Net.WebClient)..."

Expected Outcome: The inlet filter flags the dangerous string signature and halts execution, logging the incident for audit tracing.

Enterprise Considerations and Best Practices

While basic regex and string matching provide an excellent starting foundation, securing enterprise AI requires an iterative lifecycle approach. Consider implementing the following strategies to mature your pipeline architecture:

  • Advanced PII Redaction: Integrate structured libraries such as Microsoft Presidio directly into your Python Pipeline script to detect and mask social security numbers, medical identities, and names dynamically rather than blocking the prompt entirely.
  • Centralized Audit Logging: Modify the inlet function to send flagged incidents directly to your corporate Security Information and Event Management (SIEM) system like Splunk or Datadog via a secure webhook.
  • Performance Tuning: Complex regex and external API lookups add latency to user prompts. Ensure your code executes asynchronously and cache heavily queried compliance validation sets.

Conclusion

Customizing Open-WebUI Pipelines gives your enterprise the absolute control required to deploy generative AI applications safely. By building custom Guardrail filters, you can effectively minimize vulnerability to corporate espionage, data exfiltration, and malicious attacks, paving the way for a compliant and highly productive AI ecosystem.

Implementing Open-WebUI Pipelines: Building Custom Guardrail Filters to Prevent Malware and Data Leaks in Enterprise AI | DPTCloud