EVENT STREAMING ARCHITECTURE
Why Kafka KRaft Is More Than 'Kafka Without ZooKeeper'
A practical architecture conversation about metadata quorum, simpler ownership and new operational habits.
The first KRaft proposal focused on what disappeared: no separate ZooKeeper ensemble. The better architecture review focused on what became clearer: Kafka now owned its metadata lifecycle end to end.
Start with one workload | Scale into a complete platform | Open Source Freedom | Enterprise-Grade Operations.
Published | November 2025
The question behind the obvious answer
The first KRaft proposal focused on what disappeared: no separate ZooKeeper ensemble. The better architecture review focused on what became clearer: Kafka now owned its metadata lifecycle end to end.
Removing a dependency simplifies the platform, but it does not make quorum design, controller placement, backups, rolling change or recovery automatic.
| Kraft Removes A System. It Does Not Remove Operations. |
The turning point
The team redesigned the runbook around the KRaft controller quorum, broker roles and metadata evidence instead of translating old ZooKeeper steps line by line.
The team used OSS Manager to make controlled migration from ZooKeeper metadata to KRaft reviewable and repeatable. The engine still owned its native behaviour; the control plane made intent, prerequisites, execution and evidence visible to the people responsible for the service.
What had to become explicit
- The service boundary: A Kafka KRaft controller quorum replaces the external ZooKeeper dependency. Migration is planned against a supported Kafka path, rehearsed, observed, and completed without changing application topic semantics.
- The operational controls: supported migration path, metadata backup, controller quorum sizing, broker compatibility, feature level, listener validation, rollback boundary, and client regression tests.
- The capacity conversation: broker and partition count, metadata volume, controller count, fault domains, request rate, maintenance window, and rollback retention.
- The human boundary: who may observe, who may approve, who may execute, and who decides whether the application is ready.
The architecture that changed the conversation
The diagram is deliberately centred on the decision the team had to make. It is not a product inventory. It shows where authority sits, what crosses the boundary, and where a failed assumption must stop the workflow.

The practical design choices
| Lens | ZooKeeper-era concern | KRaft operating habit |
| Ownership | Two distributed systems | One Kafka lifecycle |
| Quorum | ZooKeeper ensemble | KRaft controllers |
| Recovery | Cross-system procedure | Kafka-native evidence |
| Node identity | Broker and ZooKeeper mappings | Stable broker and controller IDs |
| Rolling change | Coordinate two control planes | Controller-majority-aware sequence |
These choices are intentionally small enough to review and test. They keep the architecture tied to operating behaviour instead of allowing a visually impressive diagram to hide unclear ownership.
What OSS Manager changed - and what it did not
OSS Manager brought discovery, planning, guarded execution and normalised status into one path for Apache Kafka. It did not replace the engine's correctness model or the application's responsibility for data semantics.
- Before change, validate versions, hosts, identities, artifacts, topology and the recovery boundary without mutating the service.
- During change, persist progress, expose stop conditions and prevent a partial result from being mistaken for success.
- After change, observe controller quorum health; active controller; metadata lag; broker registration; offline and under-replicated partitions; request latency; client errors and run representative application journeys, not only process checks.
- For recovery, protect the last known-good authority and require an explicit decision before promotion, rollback or destructive cleanup.
The failure modes worth rehearsing
The test plan should make room for unsupported source version, quorum misconfiguration, metadata mismatch, broker registration failure, mixed-mode timeout, certificate error, client regression, and premature ZooKeeper removal. The purpose is not to produce a longer checklist. It is to learn whether operators can recognise the failure and choose the safe next action while the evidence is incomplete.
A human operating model
L1 operators need a plain-language answer to what is healthy, what is delayed and whether a change is in progress. Platform engineers need topology and evidence. Application owners need to know what users will experience. A useful platform connects those views without giving every person the same privileges.
| KRaft is the stronger direction because it reduces split ownership. The benefit appears only when operations evolve with the architecture. |
Where this pattern earns its place
BFSI
This pattern is relevant where teams need controlled change, explicit authority, auditable recovery and customer-journey validation. The architecture must still be adjusted for local data classification, recovery objectives, workload shape and application behaviour.
Media and OTT
This pattern is relevant where teams need regional continuity, burst traffic, low-latency serving and predictable recovery. The architecture must still be adjusted for local data classification, recovery objectives, workload shape and application behaviour.
Retail
This pattern is relevant where teams need campaign peaks, catalogue freshness, transaction continuity and rapid rollback. The architecture must still be adjusted for local data classification, recovery objectives, workload shape and application behaviour.
The lesson we would carry into the next project
KRaft is the stronger direction because it reduces split ownership. The benefit appears only when operations evolve with the architecture.
The strongest open-source platforms are not the ones with the most automation. They are the ones where automation makes ownership, risk and recovery easier for people to understand.
| Start with one workload. Prove the operating model. Then scale it into a complete platform. |
