RESILIENT STREAMING
Two Kafka Sites, One Event Story: Building Active-Active Without Building a Loop
A hybrid-cloud architecture story about local independence, selective replication and explicit event ownership.
One media platform wanted viewers in two regions to keep streaming through a site failure. Another financial platform needed local event processing without turning every WAN interruption into a global outage.
Start with one workload | Scale into a complete platform | Open Source Freedom | Enterprise-Grade Operations.
Published | March 2026
The architecture began with a constraint
One media platform wanted viewers in two regions to keep streaming through a site failure. Another financial platform needed local event processing without turning every WAN interruption into a global outage.
Copying every topic in both directions would have looked symmetrical on a diagram and behaved unpredictably in production. Loops, duplicate events, schema drift and ambiguous ownership would simply travel faster.
| Active-Active Is A Business Ownership Design. |
The turning point
Each site remained an independent Kafka service. Only approved streams crossed the boundary; topic ownership, loop prevention, offset translation and recovery tests were designed before replication was enabled.
The team used OSS Manager to make two-site Kafka active-active with explicit ownership and loop prevention reviewable and repeatable. The engine still owned its native behaviour; the control plane made intent, prerequisites, execution and evidence visible to the people responsible for the service.
What had to become explicit
- The service boundary: Each site runs an independent KRaft Kafka cluster. Approved replication flows copy selected topics in both directions with deterministic naming, loop prevention, conflict policy, and site-local producers and consumers.
- The operational controls: topic ownership, replication allow-list, loop prevention, offset translation, schema compatibility, lag objectives, duplicate handling, and site isolation tests.
- The capacity conversation: local traffic, cross-site replicated volume, retention, partitions, replication factor, WAN bandwidth and latency, replay window, and failure headroom.
- The human boundary: who may observe, who may approve, who may execute, and who decides whether the application is ready.
The architecture that changed the conversation
The diagram is deliberately centred on the decision the team had to make. It is not a product inventory. It shows where authority sits, what crosses the boundary, and where a failed assumption must stop the workflow.

The practical design choices
| Design choice | Reason | Failure test |
| Local producers | Keep site autonomy | Isolate WAN |
| Topic allow-list | Bound cost and blast radius | Pause one flow |
| Deterministic ownership | Prevent write ambiguity | Recover and fail back |
| Offset translation | Preserve consumer continuity | Move a consumer group between sites |
| Loop prevention | Stop replicated events returning | Break one direction and restore it |
These choices are intentionally small enough to review and test. They keep the architecture tied to operating behaviour instead of allowing a visually impressive diagram to hide unclear ownership.
What OSS Manager changed - and what it did not
OSS Manager brought discovery, planning, guarded execution and normalised status into one path for Apache Kafka. It did not replace the engine's correctness model or the application's responsibility for data semantics.
- Before change, validate versions, hosts, identities, artifacts, topology and the recovery boundary without mutating the service.
- During change, persist progress, expose stop conditions and prevent a partial result from being mistaken for success.
- After change, observe cluster health; replication lag; rejected records; loop-prevention drops; consumer lag; duplicate rate; recovery time; cross-site throughput and run representative application journeys, not only process checks.
- For recovery, protect the last known-good authority and require an explicit decision before promotion, rollback or destructive cleanup.
The failure modes worth rehearsing
The test plan should make room for WAN partition, dual writes to the same logical key, replication loop, schema divergence, remote lag, site outage, offset mismatch, and failback surge. The purpose is not to produce a longer checklist. It is to learn whether operators can recognise the failure and choose the safe next action while the evidence is incomplete.
A human operating model
L1 operators need a plain-language answer to what is healthy, what is delayed and whether a change is in progress. Platform engineers need topology and evidence. Application owners need to know what users will experience. A useful platform connects those views without giving every person the same privileges.
| The architecture worked because the sites could stand alone. Replication added continuity; it did not become the only thing holding the platform together. |
Where this pattern earns its place
Retail
This pattern is relevant where teams need campaign peaks, catalogue freshness, transaction continuity and rapid rollback. The architecture must still be adjusted for local data classification, recovery objectives, workload shape and application behaviour.
Telecom
This pattern is relevant where teams need high event volume, distributed operations, identity boundaries and service assurance. The architecture must still be adjusted for local data classification, recovery objectives, workload shape and application behaviour.
Gaming
This pattern is relevant where teams need latency-sensitive sessions, leaderboards, event streams and regional player experience. The architecture must still be adjusted for local data classification, recovery objectives, workload shape and application behaviour.
The lesson we would carry into the next project
The architecture worked because the sites could stand alone. Replication added continuity; it did not become the only thing holding the platform together.
The strongest open-source platforms are not the ones with the most automation. They are the ones where automation makes ownership, risk and recovery easier for people to understand.
| Start with one workload. Prove the operating model. Then scale it into a complete platform. |
