Hardening Docker on VPS Production: Building Custom Seccomp Profiles to Restrict Dangerous System Calls
Introduction to Container Security at the Kernel Level
In the modern DevOps landscape, deploying applications using Docker on Virtual Private Servers (VPS) has become the standard for businesses seeking agility and scalability. However, a common misconception persists: that container isolation is equivalent to traditional virtual machine isolation. In reality, containers share the host operating system's kernel. If a containerized application is compromised, an attacker can interact directly with the host kernel via system calls (syscalls).
By default, Docker restricts around 44 syscalls out of more than 300 available in Linux. While this default Security Compute (Seccomp) profile provides a baseline level of protection, it remains overly permissive for specialized production workloads. To achieve true hardening on a VPS production environment, security engineering teams must transition from generic configurations to custom Seccomp profiles tailored to the exact requirements of their applications.
Understanding Seccomp and the Risk of Unrestricted Syscalls
Seccomp is a security facility in the Linux kernel that allows a process to make a one-way transition into a state where it cannot make any system calls except for a very limited set. In the context of Docker, Seccomp acts as a firewall for syscalls, intercepting requests from containerized processes before they reach the host kernel.
Why is this critical for a production VPS? Consider the following risks:
- Privilege Escalation: Vulnerabilities within the Linux kernel (such as Dirty COW or various use-after-free bugs) often rely on obscure or advanced system calls to manipulate kernel memory. If your container does not need these syscalls, allowing them exposes your entire host to compromise.
- Container Breakouts: Attackers who gain execution capabilities inside a container will immediately look for ways to escape to the host. System calls related to mounting filesystems, loading kernel modules, or altering network routing are prime targets.
- Lateral Movement: A compromised container with excessive syscall privileges can probe the host VPS network configuration, potentially exposing other containers or databases residing on the same loopback interface.
Security Maxim: The principle of least privilege should not stop at file permissions or network ports; it must extend directly to the kernel interface via system calls.
The Anatomy of a Docker Seccomp Profile
Docker utilizes JSON-formatted files to define Seccomp profiles. A profile typically consists of three primary components:
- Default Action: The fallback action taken if a system call does not match any specified rules. In a hardening scenario, this is often set to
SCMP_ACT_ERRNO(block and return an error) orSCMP_ACT_ALLOW(if building a traditional blocklist). - Architectures: The target CPU architectures (e.g.,
amd64,x86_64) the profile applies to, ensuring system calls map correctly to their architecture-specific numbers. - Syscalls Array: A structured list of system calls alongside the explicit action to take (e.g.,
SCMP_ACT_ALLOW) and optional arguments to filter specific conditions.
Below is a conceptual example of a hardened Seccomp structure designed to explicitly allow only safe, foundational operations while rejecting everything else:
{
"defaultAction": "SCMP_ACT_ERRNO",
"architectures": ["SCMP_ARCH_X86_64"],
"syscalls": [
{
"names": ["read", "write", "exit", "fstat"],
"action": "SCMP_ACT_ALLOW"
}
]
}Step-by-Step Guide: Profiling Your Application's Syscall Fingerprint
Implementing a strict white-list Seccomp profile requires knowing precisely which system calls your production application needs to function normally. Blocking a required system call will result in immediate application crashes or silent runtime errors (e.g., Operation not permitted). Follow this structured approach to audit your containerized application safely:
1. Auditing with Strace or Extended BPF (eBPF)
Before deploying to production, run your containerized application in a staging environment while attached to an auditing tool. strace is highly useful for tracing system calls, though eBPF-based tools (such as Falco or Inspector Gadget) are preferred for production-like loads due to lower performance overhead.
To capture syscalls using strace during an automated integration test, you can initialize the process as follows:
strace -c -f -p This will generate a summary table listing each syscall executed by the application and its respective frequency.
2. Leveraging Docker's Audit Mode
Alternatively, you can craft a Seccomp profile that logs system calls instead of blocking them outright. By changing the defaultAction to SCMP_ACT_LOG, Docker will allow the container to execute any system call but will record every occurrence in the host's system logs (/var/log/syslog or via journalctl). Reviewing these logs after running comprehensive application tests provides an exact blueprint for your custom whitelist.
Deploying and Verifying the Custom Profile on a Production VPS
Once you have compiled the definitive list of required system calls, you can generate your production-grade JSON profile. Deploying it requires modifying your container runtime execution commands or orchestration configurations.
Executing via Docker CLI
To launch a standalone container on your VPS with the custom profile enforced, utilize the --security-opt flag, passing the path to your JSON configuration file:
docker run -d \
--name secure-app-service \
--security-opt seccomp=/path/to/custom-seccomp.json \
-p 8080:8080 \
my-production-image:v1.0Integration with Docker Compose
For maintainable production environments managed via Docker Compose, embed the security options directly within the service definitions of your docker-compose.yml file:
version: '3.8'
services:
web_application:
image: my-production-image:v1.0
ports:
- "8080:8080"
security_opt:
- seccomp:/etc/docker/seccomp/custom-seccomp.json
restart: alwaysVerification and Continuous Monitoring
To verify that the custom Seccomp profile is successfully applied, inspect the running container configuration using the Docker engine CLI:
docker inspect secure-app-service | grep SeccompEnsure that the output reflects the path to your custom profile rather than indicating the default or unconfined state. Furthermore, monitor your application logs continuously during the post-deployment phase to guarantee that no edge-case application pathways are being blocked unexpectedly.
Conclusion and Operational Best Practices
Restricting kernel system calls through custom Seccomp profiles is one of the most effective measures an organization can take to harden Docker on VPS instances. It bridges the gap between simple namespace isolation and robust container security, mitigating risks associated with zero-day kernel exploits and privilege escalation.
To maintain a high security posture over time, adhere to these operational best practices:
- Incorporate Seccomp into CI/CD pipelines: Automate syscall auditing as part of your integration testing phase to catch changing requirements before deployment.
- Never run containers as root: Combining a custom Seccomp profile with a non-root container user provides a powerful, multi-layered defense mechanism.
- Keep profiles under version control: Treat your Seccomp JSON configurations as Infrastructure as Code (IaC) to track modifications and facilitate seamless rollbacks if issues emerge.
