Back to articles
Technology Insight

Building a Lightweight CDC Pipeline: Syncing MySQL to Meilisearch with Debezium on a VPS

May 30, 2026

Introduction: The Quest for Real-Time Search Without the Overhead

In modern web applications, search functionality is no longer a luxury—it is a core user expectation. When users type a query into a search bar, they expect instantaneous, relevant results. While traditional relational databases like MySQL are exceptional for transactional integrity, they inherently struggle with complex full-text search queries, especially at scale.

To solve this, developers frequently introduce dedicated search engines. Meilisearch has emerged as a premier choice for this role: it is open-source, lightning-fast, and designed for instant, typo-tolerant search experiences. However, introducing a search engine creates a critical architectural challenge: how do you keep the search index synchronized with your primary database in real time?

Traditional approaches like cron jobs or application-level hooks introduce latency, race conditions, or tight coupling. This is where Change Data Capture (CDC) comes in. In this technical guide, we will explore how to build a highly efficient, lightweight CDC pipeline from MySQL to Meilisearch using Debezium, optimized specifically to run on a budget-friendly Virtual Private Server (VPS) without draining system resources.

Understanding the Architecture: Why Debezium and Meilisearch?

Before diving into the configuration, it is essential to understand why this specific stack offers a superior balance of performance and resource consumption on a VPS.

  • MySQL Binlog: MySQL records every data modification (INSERT, UPDATE, DELETE) into its binary log (binlog). By reading this log, we can capture changes asynchronously without impacting application performance.
  • Debezium: Debezium is a distributed, open-source platform built on top of Apache Kafka connect paradigms. For our ultra-lightweight VPS setup, we will utilize Debezium Server, a standalone distribution that eliminates the heavy dependency on a full Kafka cluster, routing events directly to target systems or lightweight brokers.
  • Meilisearch: Unlike Elasticsearch, which requires massive JVM heaps, Meilisearch is written in Rust. It features an incredibly small memory footprint and exceptional CPU efficiency, making it the perfect candidate for VPS deployment.
Architectural Note: By bypassing a full Apache Kafka cluster and leveraging Debezium Server alongside Meilisearch, we can run this entire real-time synchronization pipeline on a standard VPS with as little as 2GB to 4GB of RAM.

Step 1: Preparing the MySQL Database for CDC

To enable Debezium to read database modifications, MySQL must be configured to output row-based binary logs. Modify your MySQL configuration file (typically my.cnf or mysqld.cnf) to include the following directives:

[mysqld]
server-id         = 1
log-bin           = mysql-bin
binlog_format     = ROW
binlog_row_image  = FULL
expire_logs_days  = 7

After updating the configuration, restart your MySQL service to apply the changes. Next, create a dedicated database user with the minimum required privileges for Debezium to monitor the binlog:

CREATE USER 'debezium'@'%' IDENTIFIED BY 'YourSecurePassword';
GRANT SELECT, RELOAD, SHOW DATABASES, REPLICATION SLAVE, REPLICATION CLIENT ON *.* TO 'debezium'@'%';
FLUSH PRIVILEGES;

These permissions allow Debezium to perform an initial snapshot of the database data and continuously stream subsequent row-level changes.

Step 2: Deploying Meilisearch on Your VPS

Meilisearch can be installed natively or via Docker. For isolated and predictable resource management on a VPS, Docker is highly recommended. Run the following command to deploy Meilisearch securely:

docker run -d -p 7700:7700 \
  -v $(pwd)/meili_data:/meili_data \
  -e MEILI_ENV="production" \
  -e MEILI_MASTER_KEY="A_Very_Strong_Master_Key_For_Security" \
  --name meilisearch \
  getmeili/meilisearch:v1.x

Ensure you save the MEILI_MASTER_KEY. This token will be required by our consumer application to write synchronized data into the Meilisearch indexes.

Step 3: Configuring the Lightweight Debezium Engine

To keep our stack "ultra-lightweight," we will avoid deploying Kafka, Zookeeper, and Kafka Connect. Instead, we will use a custom consumer script paired with Debezium Embedded or route Debezium Server events to a lightweight message broker like Redis or MQTT, which Meilisearch can consume. Alternatively, we can use a small Node.js or Python daemon that acts as an HTTP webhook target for Debezium Server.

Let us configure Debezium Server to stream changes to a local, ultra-lightweight Redis stream. Create a application.properties file for Debezium:

debezium.sink.type=redis
debezium.sink.redis.address=localhost:6379
debezium.source.connector.class=io.debezium.connector.mysql.MySqlConnector
debezium.source.offset.storage.file.filename=data/offsets.dat
debezium.source.database.hostname=127.0.0.1
debezium.source.database.port=3306
debezium.source.database.user=debezium
debezium.source.database.password=YourSecurePassword
debezium.source.database.server.id=1
debezium.source.database.server.name=my-app-vps
debezium.source.database.include.list=production_db
debezium.source.table.include.list=production_db.products

This structure ensures that any change inside the products table of production_db instantly converts into a structured JSON event payload inside Redis, consuming negligible CPU cycles.

Step 4: Implementing the Meilisearch Consumer Sync Engine

With data flowing smoothly into our lightweight Redis buffer, we need a simple, event-driven script to consume these records and update Meilisearch. Below is a highly optimized Python script using standard libraries to transform Debezium CDC events into Meilisearch API updates.

import redis
import requests
import json

# Configuration
REDIS_HOST = '127.0.0.1'
REDIS_PORT = 6379
MEILI_URL = '[http://127.0.0.1:7700](http://127.0.0.1:7700)'
MEILI_KEY = 'A_Very_Strong_Master_Key_For_Security'
INDEX_NAME = 'products'

r = redis.Redis(host=REDIS_HOST, port=REDIS_PORT, decode_responses=True)
headers = {'Authorization': f'Bearer {MEILI_KEY}', 'Content-Type': 'application/json'}

print("CDC Sync Engine started successfully...")

while True:
    # Read from Redis stream
    events = r.xread({'debezium.events': '0-0'}, block=0, count=10)
    for stream, messages in events:
        for msg_id, payload in messages:
            event_data = json.loads(payload['value'])
            op = event_data.get('op') # c: create, u: update, d: delete
            
            if op in ['c', 'u']:
                # Extract the post-image data
                document = event_data['after']
                requests.post(f"{MEILI_URL}/indexes/{INDEX_NAME}/documents", json=[document], headers=headers)
            elif op == 'd':
                # Extract the pre-image ID for deletion
                doc_id = event_data['before']['id']
                requests.delete(f"{MEILI_URL}/indexes/{INDEX_NAME}/documents/{doc_id}", headers=headers)
            
            # Acknowledge or delete message from stream to save memory
            r.xdel('debezium.events', msg_id)

This daemon script runs continuously under a process supervisor like Systemd or Supervisor. It intercepts creation, update, and deletion queries instantly, transforming them into optimal payloads for Meilisearch without heavy middleware.

Optimizing Performance and Resource Management on a VPS

When running a pipeline like this on a constrained environment, continuous tuning ensures long-term stability:

  1. Memory Capping: Configure Redis with a maxmemory limit and an eviction policy (e.g., volatile-lru) to ensure it never triggers the VPS Out-Of-Memory (OOM) killer.
  2. Batching Updates: In production environments with high write volumes, modify the consumer script to accumulate events over 500ms or up to 100 documents before pushing them to Meilisearch in a single batch API call. This dramatically reduces HTTP overhead.
  3. Index Provisioning: Define your Meilisearch filterable and searchable attributes explicitly beforehand. This reduces internal compilation overhead whenever documents are synchronized.

Conclusion: Enterprise-Grade Search on a Budget

By leveraging the power of row-based binary logging with Debezium Server and combining it with the lightweight execution model of Meilisearch, we successfully built an enterprise-ready, real-time search sync pipeline. This architecture eliminates the heavy operational complexities of full Kafka clusters, making it perfectly suited for small to mid-sized applications hosted on affordable VPS setups.

With sub-second indexing latency and negligible resource overhead, your users can now enjoy instantaneous search capabilities without exploding your monthly infrastructure costs.

Building a Lightweight CDC Pipeline: Syncing MySQL to Meilisearch with Debezium on a VPS | DPTCloud