Back to articles
Technology Insight

Diskless Centralized Logging: Streaming Docker Container Logs Directly to AWS S3 and Cloudflare R2 Using Vector.dev

June 4, 2026

Introduction to Modern Log Management Challenges

In contemporary cloud-native environments, logging is both a critical operational necessity and a frequent source of infrastructure bottlenecks. Traditional logging paradigms dictate that applications write logs to local files, which are subsequently collected, rotated, and shipped by a background daemon. While this approach is straightforward, it introduces severe disk I/O contention, risks data loss during local storage depletion, and increases operational overhead when managing massive containerized deployments.

As enterprises scale their microservices using Docker and Kubernetes, the volume of telemetry data grows exponentially. Relying on local node storage for temporary log buffering creates an architectural vulnerability. A sudden spike in application traffic can lead to a deluge of logs, saturating the local disk, degrading application performance, and potentially crashing the entire host node. To mitigate these risks, forward-thinking engineering teams are pivoting toward a diskless centralized logging architecture.

The Philosophy of Diskless Centralized Logging

The core philosophy of diskless centralized logging is simple yet transformative: eliminate the local filesystem as a middleman for log data. Instead of persisting logs to a virtual or physical drive on the host machine, logs are intercepted directly from the runtime container engine stdout/stderr streams and immediately transmitted across the network to highly durable, cost-effective object storage systems.

By removing local disk writes, organizations achieve several key architectural benefits:

  • Zero Local Disk I/O Overhead: Eliminates write amplification on ephemeral host drives, preserving disk bandwidth for primary database and application workloads.
  • Enhanced Security and Compliance: Logs never reside on local disks where they could be compromised or prematurely deleted by an intruder or a system crash.
  • Substantial Cost Reductions: Storing terabytes of historical logs on high-performance block storage (like AWS EBS) is financially unsustainable. Streaming directly to object storage (like AWS S3 or Cloudflare R2) reduces storage costs by up to 90%.

Introducing Vector.dev: The High-Performance Telemetry Router

To implement a diskless logging pipeline, you need a high-performance, lightweight data router capable of processing millions of events per second with minimal memory footprint. While legacy tools like Fluentd, Logstash, and Filebeat have served the industry well, they often suffer from high CPU consumption and significant memory overhead due to their underlying runtimes (Ruby/JVM).

Enter Vector.dev, an open-source, ultra-fast telemetry agent written in Rust. Developed by Datadog, Vector is designed from the ground up for safety, speed, and resource efficiency. It acts as an end-to-end observability pipeline that can collect, transform, and route logs, metrics, and traces. Vector’s memory safety guarantees and asynchronous I/O engine make it the ideal candidate for handling high-throughput Docker log streams without introducing latency to host applications.

Architectural Overview: From Docker to Cloud Object Storage

The diskless logging pipeline constructed in this guide leverages three fundamental components acting in unison:

  1. The Source (Docker Daemon): Docker containers run with their logging driver configured to standard output. Vector connects directly to the Docker daemon via the local UNIX socket (/var/run/docker.sock), intercepting the raw log streams in real-time before they are written to standard JSON log files on the host disk.
  2. The Router (Vector.dev): Vector ingests the container logs, parses the metadata (such as container name, image, and timestamps), structures the payloads into a unified format, and optionally filters or anonymizes sensitive data using its powerful Vector Remap Language (VRL).
  3. The Sinks (AWS S3 & Cloudflare R2): Vector buffers the processed logs entirely in memory before flushing them in optimized batches directly to cloud object storage.
Note on Cloudflare R2: Because Cloudflare R2 is fully compatible with the AWS S3 API ecosystem, Vector can treat R2 as an S3-compatible sink. Cloudflare R2’s primary advantage is its zero-egress fee model, making it exceptionally economical for long-term log retention and analytical querying.

Step-by-Step Implementation Guide

Step 1: Preparing the Cloud Infrastructure

Before configuring Vector, you must provision your storage buckets and set up secure access permissions. For AWS S3, create a bucket named enterprise-docker-logs and attach an IAM policy granting s3:PutObject permissions to your Vector execution role. For Cloudflare R2, generate an API token with Edit permissions and note your unique Account ID, Access Key ID, and Secret Access Key.

Step 2: Configuring Vector (vector.yaml)

Create a centralized configuration file named vector.yaml to orchestrate the pipeline. Below is a production-ready configuration that defines the Docker source, an internal VRL transformation step, and dual sinks for AWS S3 and Cloudflare R2.


sources:
  docker_containers:
    type: "docker_logs"
    include_containers: [] # Leave empty to include all containers

transforms:
  parse_and_enrich:
    type: "remap"
    inputs:
      - "docker_containers"
    source: |
      .structured = parse_json(.message) ?? {}
      .environment = "production"
      .timestamp = .timestamp || now()

sinks:
  aws_s3_storage:
    type: "aws_s3"
    inputs:
      - "parse_and_enrich"
    bucket: "enterprise-docker-logs"
    region: "us-east-1"
    compression: "gzip"
    filename_append_uuid: true
    key_prefix: "docker-logs/year=%Y/month=%m/day=%d/"
    batch:
      max_bytes: 5242880 # 5MB batches
      timeout_secs: 30

  cloudflare_r2_storage:
    type: "aws_s3"
    inputs:
      - "parse_and_enrich"
    bucket: "production-logs-r2"
    endpoint: "https://.r2.cloudflarestorage.com"
    region: "auto"
    compression: "gzip"
    filename_append_uuid: true
    key_prefix: "docker-logs/year=%Y/month=%m/day=%d/"
    auth:
      access_key_id: "${R2_ACCESS_KEY_ID}"
      secret_access_key: "${R2_SECRET_ACCESS_KEY}"
    batch:
      max_bytes: 5242880
      timeout_secs: 30

Step 3: Deploying Vector via Docker Compose

To execute Vector without storing logs on the host, run Vector itself as a lightweight sidecar container. It requires access to the Docker socket to read other containers' logs, but its own internal buffer is explicitly restricted to system memory. version: "3.8" services: vector: image: timberio/vector:0.34.0-alpine container_name: vector_log_router volumes: - ./vector.yaml:/etc/vector/vector.yaml:ro - /var/run/docker.sock:/var/run/docker.sock:ro environment: - R2_ACCESS_KEY_ID=your_r2_key_here - R2_SECRET_ACCESS_KEY=your_r2_secret_here restart: always network_mode: "host"

Best Practices for Production Environments

While a diskless architecture removes local storage bottlenecks, relying entirely on network-based streaming requires careful configuration to guarantee stability and prevent data loss during network partitions.

1. Memory Buffer Management

Because logs are not written to disk, Vector holds active log batches in the system RAM before flushing them to cloud providers. You must ensure the host machine has adequate memory allocation. If cloud endpoints experience transient downtime, Vector’s memory buffer can fill up. Configure Vector's buffer.when_full behavior to either block (slowing down application container logging slightly to preserve data) or drop_newest (prioritizing application availability over log completeness), depending on your business compliance requirements.

2. Optimized Hive Partitioning

Notice the key_prefix configuration used in the example: docker-logs/year=%Y/month=%m/day=%d/. Organizing logs using this format structures the files cleanly in S3/R2. This structure matches standard Hive partitioning schemas, allowing distributed SQL query engines like AWS Athena or DuckDB to query billions of log lines instantly while minimizing data scanning costs.

Conclusion

Transitioning to a diskless centralized logging architecture using Vector.dev, Docker, and object storage like AWS S3 or Cloudflare R2 represents a massive paradigm shift in cloud infrastructure management. By shifting data persistence away from high-latency, expensive local block storage and into highly resilient object stores, enterprises can drastically lower total cost of ownership while eliminating a critical point of failure. Implementing Vector provides the performance guarantees of Rust, ensuring your infrastructure scales seamlessly alongside your application workloads.