HOME/BLOG/technology
technology

Kafka Active-Active: Two Sites, One Event Story

Explore Kafka active-active across two sites with selective replication, loop prevention, clear ownership and governed recovery.

AUTHOR: Chinmay Kulkarni•02 March 2026•5 MIN READ
Kafka Active-Active: Two Sites, One Event Story


RESILIENT STREAMING

Two Kafka Sites, One Event Story: Building Active-Active Without Building a Loop

A hybrid-cloud architecture story about local independence, selective replication and explicit event ownership.

One media platform wanted viewers in two regions to keep streaming through a site failure. Another financial platform needed local event processing without turning every WAN interruption into a global outage.

Start with one workload | Scale into a complete platform | Open Source Freedom | Enterprise-Grade Operations.

Published | March 2026

The architecture began with a constraint

One media platform wanted viewers in two regions to keep streaming through a site failure. Another financial platform needed local event processing without turning every WAN interruption into a global outage.

Copying every topic in both directions would have looked symmetrical on a diagram and behaved unpredictably in production. Loops, duplicate events, schema drift and ambiguous ownership would simply travel faster.

Active-Active Is A Business Ownership Design.

The turning point

Each site remained an independent Kafka service. Only approved streams crossed the boundary; topic ownership, loop prevention, offset translation and recovery tests were designed before replication was enabled.

The team used OSS Manager to make two-site Kafka active-active with explicit ownership and loop prevention reviewable and repeatable. The engine still owned its native behaviour; the control plane made intent, prerequisites, execution and evidence visible to the people responsible for the service.

What had to become explicit

  • The service boundary: Each site runs an independent KRaft Kafka cluster. Approved replication flows copy selected topics in both directions with deterministic naming, loop prevention, conflict policy, and site-local producers and consumers.
  • The operational controls: topic ownership, replication allow-list, loop prevention, offset translation, schema compatibility, lag objectives, duplicate handling, and site isolation tests.
  • The capacity conversation: local traffic, cross-site replicated volume, retention, partitions, replication factor, WAN bandwidth and latency, replay window, and failure headroom.
  • The human boundary: who may observe, who may approve, who may execute, and who decides whether the application is ready.

The architecture that changed the conversation

The diagram is deliberately centred on the decision the team had to make. It is not a product inventory. It shows where authority sits, what crosses the boundary, and where a failed assumption must stop the workflow.


The practical design choices

Design choiceReasonFailure test
Local producersKeep site autonomyIsolate WAN
Topic allow-listBound cost and blast radiusPause one flow
Deterministic ownershipPrevent write ambiguityRecover and fail back
Offset translationPreserve consumer continuityMove a consumer group between sites
Loop preventionStop replicated events returningBreak one direction and restore it

These choices are intentionally small enough to review and test. They keep the architecture tied to operating behaviour instead of allowing a visually impressive diagram to hide unclear ownership.

What OSS Manager changed - and what it did not

OSS Manager brought discovery, planning, guarded execution and normalised status into one path for Apache Kafka. It did not replace the engine's correctness model or the application's responsibility for data semantics.

  • Before change, validate versions, hosts, identities, artifacts, topology and the recovery boundary without mutating the service.
  • During change, persist progress, expose stop conditions and prevent a partial result from being mistaken for success.
  • After change, observe cluster health; replication lag; rejected records; loop-prevention drops; consumer lag; duplicate rate; recovery time; cross-site throughput and run representative application journeys, not only process checks.
  • For recovery, protect the last known-good authority and require an explicit decision before promotion, rollback or destructive cleanup.

The failure modes worth rehearsing

The test plan should make room for WAN partition, dual writes to the same logical key, replication loop, schema divergence, remote lag, site outage, offset mismatch, and failback surge. The purpose is not to produce a longer checklist. It is to learn whether operators can recognise the failure and choose the safe next action while the evidence is incomplete.

A human operating model

L1 operators need a plain-language answer to what is healthy, what is delayed and whether a change is in progress. Platform engineers need topology and evidence. Application owners need to know what users will experience. A useful platform connects those views without giving every person the same privileges.

The architecture worked because the sites could stand alone. Replication added continuity; it did not become the only thing holding the platform together.

Where this pattern earns its place

Retail

This pattern is relevant where teams need campaign peaks, catalogue freshness, transaction continuity and rapid rollback. The architecture must still be adjusted for local data classification, recovery objectives, workload shape and application behaviour.

Telecom

This pattern is relevant where teams need high event volume, distributed operations, identity boundaries and service assurance. The architecture must still be adjusted for local data classification, recovery objectives, workload shape and application behaviour.

Gaming

This pattern is relevant where teams need latency-sensitive sessions, leaderboards, event streams and regional player experience. The architecture must still be adjusted for local data classification, recovery objectives, workload shape and application behaviour.

The lesson we would carry into the next project

The architecture worked because the sites could stand alone. Replication added continuity; it did not become the only thing holding the platform together.

The strongest open-source platforms are not the ones with the most automation. They are the ones where automation makes ownership, risk and recovery easier for people to understand.

Start with one workload. Prove the operating model. Then scale it into a complete platform.