Building a DIY AI-Enhanced Security Operations Center on a VPS: Integrating Wazuh, TheHive, and Anomaly Detection for Small Networks
Introduction: The Democratization of Enterprise-Grade Security
For years, comprehensive Security Operations Centers (SOCs) were the exclusive domain of large enterprises with substantial budgets and dedicated security teams. Small businesses, startups, and even tech-savvy individuals were left with piecemeal solutions that offered limited visibility and required constant manual oversight. This security gap created significant risk, as threat actors increasingly target smaller networks precisely because of their weaker defenses.
The landscape is changing. The convergence of affordable cloud Virtual Private Servers (VPS), mature open-source security tools, and accessible artificial intelligence frameworks has made it possible to build a sophisticated, AI-enhanced SOC on a modest budget. This guide will walk you through creating a "SOC at home" by integrating Wazuh (for intrusion detection and log management), TheHive (for security incident response and case management), and custom AI models for behavioral anomaly detection.
Why a DIY, AI-Enhanced SOC Makes Sense for Small Networks
Before diving into the technical implementation, it's crucial to understand the value proposition. A traditional, manual monitoring approach is reactive and scales poorly. An AI-enhanced SOC built on a VPS offers several distinct advantages:
- Proactive Threat Detection: Move beyond signature-based detection to identify novel attacks and insider threats through behavioral analysis.
- Cost Efficiency: Leverage open-source software and a mid-tier VPS (\~$20-40/month) to achieve capabilities that would cost thousands in commercial solutions.
- Automated Triage & Response: Reduce alert fatigue by using AI to correlate events, score severity, and even initiate predefined automated responses.
- Centralized Visibility: Gain a single pane of glass for security events across your entire network, from servers and workstations to network devices and cloud services.
- Skill Development & Customization: Build a system tailored to your specific environment and threat model, gaining invaluable hands-on security engineering experience.
Architectural Blueprint: Core Components and Data Flow
The effectiveness of your SOC hinges on a clean, scalable architecture. We propose a modular design where each component has a defined role, communicating through well-established protocols.
1. The Foundation: Wazuh as the Security Data Lake
Wazuh serves as the backbone. It is an open-source platform for threat detection, integrity monitoring, incident response, and compliance. Its agents, deployed on endpoints (servers, laptops, etc.), collect a vast array of data: system logs, file integrity information, process inventories, and vulnerability detection results. This data is sent to the Wazuh manager, which normalizes and enriches it, applying a first layer of rules to generate security alerts.
2. The Brain: TheHive for Incident Orchestration
TheHive acts as the command center. It receives alerts from Wazuh (and other sources) via connectors. Each alert can be converted into a case—a container for all related information, evidence, tasks, and analyst notes. TheHive's power lies in its workflow automation, collaboration features, and integration capabilities. It's where human judgment meets automated processes.
3. The Sixth Sense: The AI Anomaly Detection Layer
This is the differentiating component. While Wazuh has basic anomaly detection, we integrate a dedicated AI model. This model, likely built with Scikit-learn, TensorFlow, or PyTorch, analyzes the stream of normalized events from Wazuh. It learns the "normal" behavior of your network—typical login times, common processes, standard network traffic patterns—and flags significant deviations. These AI-generated anomaly alerts are fed directly into TheHive as high-fidelity signals, prioritizing truly suspicious activity.
Data Flow Summary
- Collection: Wazuh agents → Wazuh Manager.
- Enrichment & Initial Detection: Wazuh Manager applies rules, generates standard alerts.
- AI Analysis: A subset of event data is streamed to the custom AI model for behavioral analysis.
- Alert Fusion: Both Wazuh rule alerts and AI anomaly alerts are sent to TheHive.
- Orchestration & Response: TheHive triages alerts, creates cases, manages workflows, and can execute automated response playbooks (e.g., isolate a host via API).
Step-by-Step Implementation Guide
This section outlines the practical steps to build the system on a Linux-based VPS (Ubuntu 22.04 LTS is recommended).
Phase 1: VPS Provisioning and Base Setup
Select a VPS provider (DigitalOcean, Linode, Vultr, AWS Lightsail) with at least 4GB RAM, 2 vCPUs, and 80GB SSD storage. Security starts here:
- Harden the SSH configuration (disable root login, use key-based auth).
- Configure a firewall (UFW) to allow only necessary ports (SSH, Wazuh/HTTPS, TheHive).
- Install Docker and Docker Compose, which will simplify the deployment of TheHive and its dependencies (Cortex for analyzers, MISP for threat intel).
Phase 2: Deploying the Wazuh Stack
The easiest method is using Wazuh's official All-in-One (AIO) deployment via Docker Compose. This single deployment includes the manager, indexer (Elasticsearch fork), and dashboard (Kibana fork). After deployment, you will access the Wazuh dashboard via HTTPS to monitor agent status and view alerts. Next, generate and install Wazuh agents on the endpoints you wish to monitor.
Phase 3: Deploying TheHive, Cortex, and MISP
Use the official Docker Compose files for TheHive 4.x. This will typically spin up three core services:
- TheHive: The main application.
- Cortex: Allows TheHive to enrich data by querying external analyzers (VirusTotal, Shodan, etc.).
- MISP (optional but recommended): A threat intelligence platform to feed IOCs (Indicators of Compromise) into your SOC.
Once running, configure the Wazuh-TheHive integration. This involves setting up a webhook in Wazuh to forward its alerts to TheHive's API and creating a corresponding application in TheHive to receive them.
Phase 4: Building and Integrating the AI Anomaly Detector
This is the most custom part of the build. A practical approach is to create a Python service that performs the following:
- Data Ingestion: Subscribe to the Wazuh manager's event stream (via its REST API or by reading from the indexer). Focus on key event types: authentication logs (SSH, Windows logins), process creation, and network connections.
- Feature Engineering: Transform raw log data into numerical features a model can understand. Examples: login attempts per hour per user, count of unique processes launched by a user, volume of outbound connections to new destinations.
- Model Training & Inference: Start with an unsupervised learning model like an Isolation Forest or One-Class SVM. Train it on a period of known-good activity to learn the baseline. The model will then score new events; high anomaly scores trigger an alert.
- Alerting: Package high-score anomalies into a formatted alert and POST it directly to TheHive's API, mimicking the Wazuh connector. This ensures all alerts are managed in a single platform.
Pro Tip: Begin with a simple model and a narrow data source (e.g., only SSH logs). It's better to have a highly accurate detector for one critical vector than a noisy, unreliable one for everything. Iteratively expand scope as you tune the model.
Operationalizing Your DIY SOC
Building the system is only half the battle. Operational excellence is key to deriving value.
Fine-Tuning and Reducing False Positives
Initially, your AI model and Wazuh rules will generate many false positives. This is normal. Use TheHive to quickly review and close false alerts. For the AI, each closed false positive is a labeled data point you can use to retrain and improve the model. Continuously curate Wazuh's decoders and rules to match your environment.
Developing Response Playbooks
Turn your SOC from a detection tool into a response engine. Within TheHive, define response playbooks for common alert types. For example, a playbook for a "brute force attack" alert could automatically:
- Enrich the attacking IP address via Cortex (AbuseIPDB, VirusTotal).
- Create a task for an analyst to review.
- If the confidence is high, automatically add the IP to the edge firewall block list via an API call.
Maintenance and Scaling Considerations
Monitor the health of your VPS (CPU, memory, disk I/O). Wazuh's indexer can be resource-intensive. Implement log rotation and retention policies. As your network grows, consider splitting components across multiple VPS instances (e.g., separate Wazuh indexer and TheHive server). Regularly update all components to incorporate security patches and new features.
Conclusion: Taking Control of Your Network Security Posture
Constructing an AI-enhanced SOC on a VPS is a challenging but immensely rewarding project. It moves you from a passive, vulnerable position to one of active defense and deep understanding. You are not just installing software; you are engineering a system that learns the unique rhythm of your network and stands guard against deviations.
The integration of Wazuh, TheHive, and custom AI creates a synergistic effect greater than the sum of its parts: comprehensive data collection, efficient incident management, and intelligent, proactive detection. This approach proves that advanced, automated security operations are no longer gated by budget, but by knowledge and initiative. Start small, focus on a critical asset, and iteratively build your capability. The peace of mind and defensive strength gained will be well worth the investment.
