Back to articles
Technology Insight

Architecting High-Performance Hybrid Caching: Two-Tier Cache with Redis and Go

August 14, 2026

Architecting High-Performance Hybrid Caching: Two-Tier Cache with Redis and Go

Introduction

In modern, high-throughput microservice architectures, sub-millisecond API response times are no longer a luxury—they are a core requirement. While distributed caching solutions like Redis provide excellent shared storage and scalability, fetching data over the network still introduces latency (typically 1ms to 5ms depending on network topology) and consumes valuable network bandwidth under extreme loads.

To achieve true ultra-low latency (sub-microsecond) reads, enterprise architectures must employ a hybrid caching strategy (also known as a two-tier cache). This pattern leverages a highly efficient local in-memory cache (L1) residing directly within the application process, coupled with a centralized distributed cache (L2) like Redis. This article details the architecture, data consistency trade-offs, and a production-ready implementation of a hybrid cache in Go, complete with real-time Pub/Sub cache invalidation.

Core Value Proposition

By layering your caching strategy, you combine the absolute speed of local memory with the global state awareness of a distributed store:

  • Near-Zero Read Latency: L1 cache hits avoid network serialization and deserialization, serving requests in nanoseconds.

  • Reduced Redis Bottlenecks: Offloading frequent reads to L1 minimizes CPU consumption, network saturation, and connection pooling pressures on your Redis cluster.

  • Resiliency Against Partitioning: If the Redis cluster experiences a brief network partition or failover, the application continues to serve reads from the L1 cache.

  • Bandwidth Conservation: Greatly reduces internal cloud data transfer costs in high-volume microservice environments.

Architecture & System Design

Implementing a hybrid cache introduces a critical architectural challenge: cache consistency. When Instance A updates a key in the database and updates L2, other instances (Instance B, C) might still have the old value cached in their respective L1 memory.

To solve this, we implement a Pub/Sub Cache Invalidation mechanism. When a mutation occurs, the mutating instance invalidates the key in L2 and publishes an invalidation event to a shared Redis channel. All other application instances subscribe to this channel and evict the corresponding key from their local L1 caches immediately.

+-------------------------------------------------------------+ | Client Request | +-------------------------------------------------------------+ | v +-------------------------+ | L1 Local Cache | (Hit: Sub-microsecond return) +-------------------------+ / \ (Miss) / \ (Invalidate Pub/Sub Message) v v +-------------------------+ +-------------------------+ | L2 Redis Cache | <---> | Redis Pub/Sub Channel | +-------------------------+ +-------------------------+ | | (Miss) / | (Broadcast Eviction) v v +-------------------------+ +-------------------------+ | Database / API | | Instance B / C (L1) | +-------------------------+ +-------------------------+

1. The L1 Layer (Local Memory)

Designed using a thread-safe, memory-bound cache. In production, this should be backed by an LRU (Least Recently Used) cache or a concurrent map with a strict maximum size constraint to prevent Out-Of-Memory (OOM) failures.

2. The L2 Layer (Distributed Memory)

An external Redis cluster storing serialized JSON or Protocol Buffers. It acts as the source of truth for caching and handles cluster-wide cache invalidation orchestration via Redis Pub/Sub.

3. The Invalidation Broker

A real-time message bus that broadcasts eviction commands to all running nodes immediately upon key modification.

Technical Deep Dive & Implementation

Below is a production-grade, thread-safe implementation of a hybrid cache in Go using the official go-redis/v9 client. It includes an in-memory map protected by an RWMutex as L1, and implements background subscription handlers to listen for invalidation commands.

package main
import (
    "context"
    "encoding/json"
    "fmt"
    "log"
    "sync"
    "time"
"github.com/redis/go-redis/v9"

)

const invalidationChannel = "cache:invalidation:channel"

type HybridCache struct { mu sync.RWMutex l1 map[string]cacheItem rdb *redis.Client instanceID string }

type cacheItem struct { value []byte expiration time.Time }

type InvalidationMsg struct { SenderID string json:"sender_id" Key string json:"key" }

func NewHybridCache(rdb redis.Client, instanceID string) HybridCache { hc := &HybridCache{ l1: make(map[string]cacheItem), rdb: rdb, instanceID: instanceID, } go hc.listenForInvalidations(context.Background()) return hc }

func (hc *HybridCache) Get(ctx context.Context, key string) ([]byte, error) { // 1. Try L1 Cache hc.mu.RLock() item, exists := hc.l1[key] hc.mu.RUnlock()

if exists && time.Now().Before(item.expiration) {
    return item.value, nil
}

// 2. Try L2 Cache (Redis)
valStr, err := hc.rdb.Get(ctx, key).Result()
if err == redis.Nil {
    return nil, fmt.Errorf("cache miss")
} else if err != nil {
    return nil, err
}

valBytes := []byte(valStr)

// 3. Set L1 with a short local TTL to prevent stale state piling up
hc.mu.Lock()
hc.l1[key] = cacheItem{
    value:      valBytes,
    expiration: time.Now().Add(5 * time.Minute),
}
hc.mu.Unlock()

return valBytes, nil

}

func (hc *HybridCache) Set(ctx context.Context, key string, value []byte, ttl time.Duration) error { // 1. Set L2 Cache (Redis) err := hc.rdb.Set(ctx, key, value, ttl).Err() if err != nil { return err }

// 2. Set L1 Cache
hc.mu.Lock()
hc.l1[key] = cacheItem{
    value:      value,
    expiration: time.Now().Add(ttl),
}
hc.mu.Unlock()

// 3. Broadcast invalidation to all other instances
msg := InvalidationMsg{
    SenderID: hc.instanceID,
    Key:      key,
}
msgBytes, _ := json.Marshal(msg)
return hc.rdb.Publish(ctx, invalidationChannel, msgBytes).Err()

}

func (hc *HybridCache) listenForInvalidations(ctx context.Context) { pubsub := hc.rdb.Subscribe(ctx, invalidationChannel) defer pubsub.Close()

ch := pubsub.Channel()
for msg := range ch {
    var payload InvalidationMsg
    if err := json.Unmarshal([]byte(msg.Payload), &payload); err != nil {
        continue
    }

    // Evict key from L1 if not triggered by ourselves
    if payload.SenderID != hc.instanceID {
        hc.mu.Lock()
        delete(hc.l1, payload.Key)
        hc.mu.Unlock()
        log.Printf("[Node: %s] Evicted key: %s due to remote update", hc.instanceID, payload.Key)
    }
}

}

Enterprise Security Hardening & Best Practices

To operate a multi-tier cache reliably at scale, several design safety nets must be applied:

  • Memory Bounds (L1 Eviction): Never use an unbounded map in Go as your L1 cache. Under memory pressure, this will lead to container OOM kills. Use a proven, concurrent, lock-free LRU cache like hashicorp/golang-lru/v2 to enforce a hard maximum item count.

  • Invalidation Storm Mitigation: If a highly popular key (hot key) is mutated frequently, invalidation broadcasts can overwhelm the app nodes, causing excessive CPU utilization during deserialization and Lock/Unlock cycles. Consider implementing a slight coalescing/throttling delay or local de-duplication for invalidation events.

  • Secure Redis Transports: Enable TLS (Transport Layer Security) for Redis connections (rediss:// schema) to guarantee that cache data and invalidation payloads traversing internal clouds are fully encrypted in transit.

  • Safe Backups (Circuit Breaking): If the Redis connection drops, ensure the hybrid cache degrades gracefully, serving exclusively from L1 if keys are still warm, or bypassing caches to access the database directly using a circuit breaker (e.g., sony/gobreaker).

Key Takeaways & Conclusion

  • Sub-millisecond Performance: Combining L1 local caching with L2 distributed caching minimizes external database and network calls, yielding ultra-high-throughput capabilities.

  • Strong Eventual Consistency: Using Redis Pub/Sub ensures that mutations to cached values are broadcast cluster-wide within milliseconds, mitigating split-brain caching issues.

  • Safe Operations: Protecting local in-memory systems from OOM failures with strict LRU limitations is mandatory for resilient cloud-native application patterns.