Oct 1, 2026

Exorcising Ghost Apache Kafka Controllers

How We Fixed Apache Kafka’s Controller Unregistration Deadlock

Olena Babenko |

RSS Feed

Staff Software Engineer

Davide Armand |

RSS Feed

Senior Software Engineer

TL;DR

  • If you experience a problem with the same symptoms–patch your Kafka with Aiven’s fix
  • Fixing design flaws in large projects like Apache Kafka is extremely difficult.
  • When you write a fix, a migration from a “patched” version to a “fix” version is an important part of a fix.

Have you ever set out to push an allegedly 15-minute fix that ended up taking multiple weeks, without even a guarantee that it would actually work? It’s happened to the best of us during our professional journeys, but this time, the story is about Apache Kafka.

Problem

It started with a simple issue we received from our customer: a freshly updated Kafka cluster refusing to enable a new Share Group feature. For some reason, a cluster leader was resisting this change. What went wrong? A misconfiguration? A misunderstanding? Stale cluster state saved somewhere?A deep dive into the cluster metadata showed unexpected results: a removed controller, that was supposed to be gone for some time already, was still considered a part of the cluster. Even though this "ghost" controller no longer existed, it was still trapped in the cluster's metadata. Because its state and features were permanently frozen in time, it was blocking the cluster from enabling the new Share Group features.This is really bad news for anyone having this problem: how do you update a machine that is no longer with us?

Solution

This exact problem was reported by Roland Sommer early this year . Around the same time, the Apache Kafka community had been designing a proper solution which would support“ unregistering” controllers. This new design allows you to permanently "remove" a controller using a built-in CLI tool. Under the hood, this command sends a new "unregistering" metadata event to the cluster. After that, a removed controller should not be taken into account during decision-making. Simple, right? Problem solved!

Well... not entirely. The problem was first introduced in Kafka version 3.9, but had the biggest impact on versions 4.0 through 4.3. Unfortunately, the described fix is only available in version 4.4. In total, there are four Kafka versions impacted, meaning any cluster created or updated within the last 1.5 years is potentially impacted! This left thousands of Aiven customers stranded with ghost controllers blocking their upgrades. We had to engineer a silver-bullet patch to fix the problem for everyone simultaneously.

Another Solution

Patching an existing, running Kafka version requires a slightly different approach because we are boxed in by a few strict limitations:

  1. The Deadline: for version 4.0, the Community End of Life (EOL) was July 2026, after which there could be no official bug fixes, only unofficial patches.
  2. The Metadata Ban: you are not allowed to change the metadata version meaning in our case we could not simply introduce a new event that says a controller is “unregistered”.
  3. The Migration Mandate: the solution has to migrate gracefully into the proposed solution in Kafka 4.4

The good news is that in patches you don’t have to be perfect; you just need to find a way to make things work. The Deadline limitation is also a huge advantage: because there was no realistic path to an official upstream fix anymore, we could apply the change to our own fork instead, where a small internal review is enough—no need for broad community sign-off.

The Metadata Ban meant introducing a new metadata event was completely off the table. But we could still use what we already had to tag the controller as a "ghost” by introducing an “decommissioned controller” feature. If a controller supports it, we simply ignore that node while making decisions. Conveniently, the current cluster leader can flag nodes with this status, bypassing the need to access offline machines.

Since we can change the problematic controller's metadata, why not just hack it so the ghost node fakes support for Share Groups? (This was the original problem the customer faced) While possible, this approach doesn't scale since you'd have to repeat the procedure for every feature and missing node. It also conflicts with the Migration Mandate. Our goal wasn't to simply hide the problem. We wanted to tag the removed controllers as "decommissioned" so we could track them and properly "unregister" them later.

Catch 22

The fix looked obvious once Kafka 4.4 arrived: unregister the controllers that no longer existed. The new `unregister-controller` command would append an `UnregisterControllerRecord` and remove their registrations from the metadata image. Problem solved.

Except the command refused to run.

KIP-1312 is deliberately protected by a metadata-version gate. Controller unregistration becomes available only after the cluster reaches `metadata.version` `4.4-IV2`. That is sensible for a new metadata record, but it created a trap for clusters carrying old registrations. The stopped controllers were no longer voters and no longer had processes, but their old `RegisterControllerRecord` entries still advertised the capabilities of their former Kafka version. Feature validation still considered those entries when deciding whether the cluster could move to 4.4. Ironically the very controllers that needed to be unregistered were exactly what stood in the way of the unregister-controller feature.

The two operations were waiting for each other:

We found the deadlock by trying both possible orders. Unregistering the the old controllers first failed with “The current MetadataVersion is too old to support controller unregistration.” Trying to finalize the metadata version first failed because ghost controllers are advertising themself as older than `4.4-IV2`.

The way out was to separate “decommissioning” from “unregistering”. The pre-4.4 Aiven patch uses the existing `RegisterControllerRecord` and adds a marker feature named `__decommissioned_controller`. This does not create a new topic or event, and does not remove the registration. It records that a stopped, non-voter controller is permanently gone. Patched controllers skip that marked registration during feature and metadata-version validation.

That gave us a bridge to the proper fix:

The difficult part was making the bridge survive the version boundary. A marker written by a 4.0-4.3 controller would be useless if the 4.4 controller did not understand it. The 4.4 forward port therefore kept the marker filter while adopting Kafka’s native physical-unregistration record. The temporary state created by the patch could now be understood by the next version and eventually cleaned up.

Solving this circular problem–where we can't unregister a controller until we update the cluster, and can't update the cluster without unregistering the controller–was the trickiest part of this fix.

Summary

When "ghost" controllers trapped in older Kafka metadata created a Catch-22 that blocked crucial cluster upgrades, we had to engineer a custom workaround. By creatively tagging these offline nodes as "decommissioned" without violating strict metadata rules, we successfully broke the deadlock. This custom patch served as a vital bridge, allowing thousands of stranded clusters to seamlessly upgrade and adopt the official upstream fix.

If you found this article because you need to fix this exact issue in your Kafka cluster, here are the relevant patches you can apply on top of your Apache Kafka GitHub fork:

If you're using the Aiven platform, the fix will be rolled out to you automatically in the latest maintenance update.