OPEN SOURCE MIGRATION
Redis to Valkey: The Migration Was Decided at the Client Boundary
A compatibility-first journey that protected application behaviour, data authority and rollback.
The team began with a reassuring fact: the applications already spoke the protocol. Then they listed the behaviours around that protocol: modules, Lua scripts, ACLs, persistence, failover discovery and client retry assumptions.
Start with one workload | Scale into a complete platform | Open Source Freedom | Enterprise-Grade Operations.
Published | October 2025
The work before the maintenance window
The team began with a reassuring fact: the applications already spoke the protocol. Then they listed the behaviours around that protocol: modules, Lua scripts, ACLs, persistence, failover discovery and client retry assumptions.
A binary replacement could pass a ping test and still fail the business workload. The risk lived in the edges where clients, modules and operational conventions met the engine.
| Compatible Commands Are Not A Complete Migration Plan. |
The turning point
Migration became a sequence of evidence: inventory actual use, validate compatibility, build the target from approved artifacts, synchronise data, shadow traffic, cut over, and preserve the return path.
The team used OSS Manager to make compatibility-led Redis-to-Valkey migration reviewable and repeatable. The engine still owned its native behaviour; the control plane made intent, prerequisites, execution and evidence visible to the people responsible for the service.
What had to become explicit
- The service boundary: A target Valkey topology is built beside the existing Redis service, populated through an approved online or offline method, validated, and cut over through a reversible endpoint change.
- The operational controls: command and module compatibility, data and TTL checks, ACL mapping, replication or export proof, dual-run policy, cutover gate, client test, and rollback endpoint.
- The capacity conversation: source data and working set, write rate, key churn, migration bandwidth, target persistence, validation sample, dual-run duration, and cutover surge.
- The human boundary: who may observe, who may approve, who may execute, and who decides whether the application is ready.
The architecture that changed the conversation
The diagram is deliberately centred on the decision the team had to make. It is not a product inventory. It shows where authority sits, what crosses the boundary, and where a failed assumption must stop the workflow.

The practical design choices
- Commands and scripts. Test: Replay representative calls. Rollback trigger: Semantic mismatch.
- Data and TTL. Test: Compare counts and expiry. Rollback trigger: Divergence.
- Clients. Test: Reconnect and failover test. Rollback trigger: Error or latency regression.
- Persistence. Test: Compare AOF and snapshot behaviour. Rollback trigger: Durability or restart mismatch.
- Modules. Test: Exercise every production dependency. Rollback trigger: Unsupported feature or output difference.
These choices are intentionally small enough to review and test. They keep the architecture tied to operating behaviour instead of allowing a visually impressive diagram to hide unclear ownership.
What OSS Manager changed - and what it did not
OSS Manager brought discovery, planning, guarded execution and normalised status into one path for Redis to Valkey. It did not replace the engine's correctness model or the application's responsibility for data semantics.
- Before change, validate versions, hosts, identities, artifacts, topology and the recovery boundary without mutating the service.
- During change, persist progress, expose stop conditions and prevent a partial result from being mistaken for success.
- After change, observe key count; TTL distribution; replication lag; command errors; memory; latency; rejected clients; data difference; rollback readiness and run representative application journeys, not only process checks.
- For recovery, protect the last known-good authority and require an explicit decision before promotion, rollback or destructive cleanup.
The failure modes worth rehearsing
The test plan should make room for unsupported module, script incompatibility, TTL drift, replication break, ACL mismatch, client-library assumption, target memory pressure, and post-cutover regression. The purpose is not to produce a longer checklist. It is to learn whether operators can recognise the failure and choose the safe next action while the evidence is incomplete.
A human operating model
L1 operators need a plain-language answer to what is healthy, what is delayed and whether a change is in progress. Platform engineers need topology and evidence. Application owners need to know what users will experience. A useful platform connects those views without giving every person the same privileges.
| The protocol made the move possible. The evidence made it safe. |
Where this pattern earns its place
Media and OTT
This pattern is relevant where teams need regional continuity, burst traffic, low-latency serving and predictable recovery. The architecture must still be adjusted for local data classification, recovery objectives, workload shape and application behaviour.
Retail
This pattern is relevant where teams need campaign peaks, catalogue freshness, transaction continuity and rapid rollback. The architecture must still be adjusted for local data classification, recovery objectives, workload shape and application behaviour.
Telecom
This pattern is relevant where teams need high event volume, distributed operations, identity boundaries and service assurance. The architecture must still be adjusted for local data classification, recovery objectives, workload shape and application behaviour.
The lesson we would carry into the next project
The protocol made the move possible. The evidence made it safe.
The strongest open-source platforms are not the ones with the most automation. They are the ones where automation makes ownership, risk and recovery easier for people to understand.
| Start with one workload. Prove the operating model. Then scale it into a complete platform. |
