REAL-TIME APPLICATIONS
Building a Redis Cache from PostgreSQL Without Teaching Every Service to Dual-Write
An event-driven cache story using PostgreSQL change data, Kafka Connect and explicit replay semantics.
A product page needed single-digit-millisecond reads, but the source of truth remained PostgreSQL. The first proposal added a cache write beside every database transaction. It also added a new failure mode to every application path.
Start with one workload | Scale into a complete platform | Open Source Freedom | Enterprise-Grade Operations.
Published | August 2025
One customer journey, several systems
A product page needed single-digit-millisecond reads, but the source of truth remained PostgreSQL. The first proposal added a cache write beside every database transaction. It also added a new failure mode to every application path.
Dual writes make application code responsible for ordering, retries and repair. When one side succeeds and the other fails, the cache quietly becomes a second, weaker database.
| The Cache Became A Projection, Not A Second Database. |
The turning point
The team projected committed PostgreSQL changes through CDC into Kafka and a managed sink. Redis keys became rebuildable views with explicit transformation, idempotency and lag signals.
The team used OSS Manager to make CDC-driven cache projection with replay and reconciliation reviewable and repeatable. The engine still owned its native behaviour; the control plane made intent, prerequisites, execution and evidence visible to the people responsible for the service.
What had to become explicit
- The service boundary: Debezium captures approved PostgreSQL changes into Kafka. A managed Kafka Connect sink projects idempotent cache records into Redis, while PostgreSQL remains the system of record and a reconciliation job detects drift.
- The operational controls: publication scope, replication slot retention, schema contract, key design, idempotent sink, delete and TTL semantics, replay policy, DLQ, backpressure, and reconciliation.
- The capacity conversation: database change rate, event size, Kafka partitions, connector tasks, transformation cost, Redis working set and TTLs, replay burst, and recovery bandwidth.
- The human boundary: who may observe, who may approve, who may execute, and who decides whether the application is ready.
The architecture that changed the conversation
The diagram is deliberately centred on the decision the team had to make. It is not a product inventory. It shows where authority sits, what crosses the boundary, and where a failed assumption must stop the workflow.

The practical design choices
| Concern | Design answer | Operational signal |
| Ordering | Partition by stable business key | Source and sink lag |
| Duplicates | Idempotent key upsert | Retry and DLQ rate |
| Repair | Replay from durable log | Rebuild duration |
| Schema | Version change events explicitly | Compatibility and sink rejection rate |
| Freshness | Measure source-to-cache delay | P95 projection age and serving latency |
These choices are intentionally small enough to review and test. They keep the architecture tied to operating behaviour instead of allowing a visually impressive diagram to hide unclear ownership.
What OSS Manager changed - and what it did not
OSS Manager brought discovery, planning, guarded execution and normalised status into one path for PostgreSQL, Kafka Connect, and Redis. It did not replace the engine's correctness model or the application's responsibility for data semantics.
- Before change, validate versions, hosts, identities, artifacts, topology and the recovery boundary without mutating the service.
- During change, persist progress, expose stop conditions and prevent a partial result from being mistaken for success.
- After change, observe PostgreSQL slot lag; Kafka consumer lag; connector task state; records and errors; DLQ volume; Redis latency and memory; cache freshness; reconciliation difference and run representative application journeys, not only process checks.
- For recovery, protect the last known-good authority and require an explicit decision before promotion, rollback or destructive cleanup.
The failure modes worth rehearsing
The test plan should make room for slot growth, schema-breaking change, duplicate event, out-of-order update, poisoned record, sink backpressure, Redis outage, stale cache, and replay storm. The purpose is not to produce a longer checklist. It is to learn whether operators can recognise the failure and choose the safe next action while the evidence is incomplete.
A human operating model
L1 operators need a plain-language answer to what is healthy, what is delayed and whether a change is in progress. Platform engineers need topology and evidence. Application owners need to know what users will experience. A useful platform connects those views without giving every person the same privileges.
| The cache became easier to trust precisely because it was allowed to be disposable and reproducible. |
Where this pattern earns its place
BFSI
This pattern is relevant where teams need controlled change, explicit authority, auditable recovery and customer-journey validation. The architecture must still be adjusted for local data classification, recovery objectives, workload shape and application behaviour.
Media and OTT
This pattern is relevant where teams need regional continuity, burst traffic, low-latency serving and predictable recovery. The architecture must still be adjusted for local data classification, recovery objectives, workload shape and application behaviour.
Retail
This pattern is relevant where teams need campaign peaks, catalogue freshness, transaction continuity and rapid rollback. The architecture must still be adjusted for local data classification, recovery objectives, workload shape and application behaviour.
The lesson we would carry into the next project
The cache became easier to trust precisely because it was allowed to be disposable and reproducible.
The strongest open-source platforms are not the ones with the most automation. They are the ones where automation makes ownership, risk and recovery easier for people to understand.
| Start with one workload. Prove the operating model. Then scale it into a complete platform. |
