Apache Kafka's documentation and operational lore for KRaft are clear on one point: you shall never lose controller quorum. It is good advice, because a Raft quorum that can no longer elect a leader cannot commit metadata, and a Kafka cluster without a functioning controller cannot create topics, change ACLs, elect partition leaders, or otherwise make progress on its control plane.
It is also advice that eventually fails someone: disks die in correlated ways, an operator recycles the wrong set of nodes, or an availability zone outage lines up with an already-degraded voter. When that happens on a dynamic KRaft quorum, the cluster does not degrade gracefully into a documented recovery path. It stops, and with stock Kafka tooling, it stays stopped indefinitely.
This scenario always means data loss, and there is typically no way to recover the data that was lost, only availability. At Aiven we treat quorum loss as an exceedingly rare but planned-for disaster: restore the control plane from the best surviving metadata log, bring the cluster back online, and be explicit that recent state on the lost disks is gone.
Most nodes do not carry quorum
On Aiven Kafka, not every service node is a controller voter. Brokers and controllers share nodes in our topology, but votership is a smaller set, typically three voters and sometimes five on larger deployments, placed across availability zones. Losing a whole AZ is usually survivable, whereas losing quorum means losing a majority of voters, which on a three-voter set is two controllers, and need not imply that two thirds of the service nodes are gone.
When quorum is lost, the KRaft metadata log stops advancing. The active controller normally keeps that log moving; the high watermark is expected to increase continuously as a majority of voters acknowledge progress. When no majority can acknowledge writes, the high watermark freezes and we alert an on-call operator. Metadata operations stay broken for as long as that condition lasts.
The same scenario will in many cases also mean partition data loss, depending on the number of brokers in the cluster, the replication factor, and whether the last appended message reached a surviving replica. But for this blog post, let's focus on quorum and the controller cluster.
Membership changes need the quorum they define
KIP-853 moved the set of voters out of static configuration and into the metadata log itself, as VotersRecord control records. Membership can now be changed with online operations, and observers learn the current voter set by replicating the same log the controllers write.
That design also creates a dependency loop the moment majority is gone. To demote the lost voters you must commit a membership change to the Raft log. To commit a membership change you need quorum. To have quorum you must remove the lost voters. Stock AddRaftVoter / RemoveRaftVoter RPCs correctly refuse to help here; they are not buggy, they are upholding Raft.
Static quorums had a blunt escape hatch — edit controller.quorum.voters and restart — but dynamic quorum does not. The log is the single source of truth, and if the only surviving copies still say that the missing nodes are voters, the surviving processes will keep waiting for them.
The same property that traps you here, the inability to change quorum voters without already having quorum, is also what makes ordinary Kafka operations stable: it provides strong guarantees that keep automated provisioning safe.
We can restore availability, but not integrity
Controller metadata stores topic and partition state, ISR and ELR information, leadership epochs, and more. If the records that described a topic existed only on disks that are gone, that topic may be gone with them. If a broker had already observed a later metadata state than the controller you revive, a naive recovery can make that broker see leadership epochs move backwards, a situation Kafka brokers cannot gracefully handle, for consistency reasons.
Losing quorum nodes therefore means, in the general case, that the cluster cannot be restored consistently. For the data and topics that are still there, though, availability can still be restored.
That work is closer to restoring a Postgres primary from a replica than to undoing a failure: some recent committed work may be missing, so you choose the surviving state that minimizes how far backwards the rest of the system is asked to go. How much is lost depends on how far the surviving nodes were lagging behind the quorum. Lost data does not come back, but partitions that survived the event can be brought online again so workloads can read and write.
Forcing a single-voter quorum
We restore availability by first inspecting every surviving node, controllers and brokers, for its on-disk copy of the __cluster_metadata log, and picking the longest one. Both roles replicate that log, and a broker observer may hold a higher log end offset than any surviving controller. Preferring that copy avoids using a metadata timeline shorter than what a surviving broker has already seen.
Kafka brokers do not gracefully handle a decrease in leadership epoch, so the recovered log must contain every metadata message those brokers have already observed. Picking the longest surviving copy, including the brokers' own copies, is how we keep them from seeing an inconsistent log.
Then, on a chosen controller, with the controller processes stopped, we bypass Raft membership rules: forcibly append VotersRecord entries to its local copy of the log, demoting every other voter until the recovered controller is the only member of the quorum. That is an offline edit of the Raft log, in the same family as formatting storage with an initial voter set, applied instead to an existing log that could no longer make progress on its own.
When that controller starts again, because it is now the only voter, it considers itself leader. Surviving nodes replicate the appended membership records, learn that they are no longer voters, and accept the continuation of the log beyond that point. Standard platform automation then notices that the cluster has fewer voters than the topology calls for, and adds replacements through ordinary online membership operations.
That choice of log is also a choice about what we are willing to discard. We refuse to recover from any copy that would move surviving participants backwards in time: if a broker has already observed a metadata record, bringing the cluster up from a shorter log would ask it to un-see that record, including leadership epochs it cannot rewind. So we take the longest surviving copy, and in doing so we accept that everything after that offset, whatever lived only on the disks that are gone, will never come back. The recovery procedure restores the longest surviving timeline.
Put together, the algorithm for restoring availability, and for cementing whatever data loss the incident already caused, is:
- Iterate to the end of the Raft log on each surviving broker and controller, and make note of the offset of its last message.
- If a broker has the longest log, copy it to any controller's data directory.
- If a controller has the longest, select that one to restart the quorum on.
- Stop the controller processes so nothing races the offline edit.
- Append
VotersRecordentries to the chosen controller's log, demoting other voters one at a time, until that controller is the only remaining voter. - Start only that controller. As the sole voter it elects itself leader and can commit again.
- Start the remaining surviving nodes and let them replicate the new membership records.
- Provision replacement nodes and promote new voters to restore the target quorum size.
Related work is brewing upstream
We do all this with bespoke tooling, because upstream Apache Kafka still has no supported way to recover from lost voter nodes. There has been some tangential work on the dev list recently, though: KIP-1347.
The KIP proposes a --override-voters flag for kafka-storage.sh format. Its primary case is not missing controllers, but stale ones: voter endpoints persisted in the metadata log that no longer resolve after a DNS or Kubernetes identity change. Controllers and brokers still believe the old hostnames are authoritative, UpdateVoter needs quorum to fix that, and quorum cannot form because nobody can reach anybody. It is the same circular dependency class as majority loss, in that you need an offline edit of what the log says about voters, but a different broken invariant: lost nodes leaving you with a quorum that can never form, versus stale DNS leaving you with the right membership but the wrong addresses.
The approach therefore does not cover last-standing recovery: when voters are gone, the problem is not how to reach them but that they remain in the set, and rewriting hostnames cannot remove them. Standing up replacement hardware under a lost voter's name fails for a deeper reason: Kafka ties each voter to a directory ID checked against the on-disk metadata store, so fresh storage is a new replica identity rather than a rename of the old one. The membership record itself has to change, demoting the lost voters, which is a different operation from overriding endpoints.
Separately, the proposed mechanism rewrites history. Rather than appending a new VotersRecord after the log end, an alternative considered in the KIP and what our recovery does, --override-voters builds a snapshot at the local log end whose voter endpoints differ from what Raft committed at that offset, so the snapshot disagrees with the log it claims to summarize.
Regardless of which design upstream eventually ships, we are glad this topic is getting attention. Dynamic quorum made membership changes safer in the common case, and it also created failure modes that stock tooling still cannot unwind. Operators and managed platforms will keep meeting those incidents; having a shared, reviewed way to talk about them, and eventually to recover from them, is progress even while the mechanism is still under discussion.
When prevention fails
Losing quorum is rare if voters are spread across failure domains and if node replacement never removes a majority at once, so prevention remains the main design. When it happens anyway, we page an operator rather than letting automation rewrite Raft history on its own. The recovery destroys any log suffix that lived only on the lost disks, and it is only appropriate once it is fully verified that those voters cannot be recovered.
Getting the control plane online again is necessary, but it is not the same as undoing the incident: recent produce data and recent metadata changes may still be irrecoverable, and the operator who runs this recovery is choosing availability over a dark cluster, which, with stock Kafka tooling, is otherwise where the service stays.
Raft requires majority for a reason, and when that majority is gone there is no honest way to pretend consensus still holds. What remains is to pick the best surviving history you have, admit what you are discarding with that choice, and start serving the workload again.

