Sep 23, 2026

Know Your data, Trust Your AI: Aiven DataHub is now GA

The fully managed, open-source data catalog that gives your teams and your AI agents the context to find, understand and trust data - with unlimited users and no per-seat licensing. Now production-ready.

Stan Dmitriev |

RSS Feed

Product Director, Aiven Context

Ask a simple question "who are our most profitable customers?" and things fall apart. The data lives in six systems, nobody agrees which table is canonical, the column called profit is actually revenue, and the business rules that matter live in someone's head or a Confluence page nobody's touched since 2023.

Now point an AI agent at that same mess. An agent without context is a brilliant intern on their first day: eager, capable, and completely clueless about your organization so it makes confident, expensive mistakes. It picks the wrong table from the wrong system, calls revenue "profit," and cheerfully exposes PII. Very confident, and very wrong.

The problem isn't your model. It's context.

Which is why today we're announcing the general availability of Aiven DataHub, a fully managed, open-source data catalog that gives people and agents the context to find, understand and trust data across every system, on Aiven and beyond, with unlimited users and no per-seat licensing, so it's a catalog the whole organization and every agent can actually use, not one rationed by seat count.

Here's what it looked like for teams who put it to work during Limited Availability.

Dojo: From "who knows where this data is?" to answers in seconds

Dojo is a major UK payments provider, and it's building AI into how the business runs; agents that pull reports and handle day-to-day operations. In payments, that AI has to be right. But Dojo's data was spread across Google Cloud Storage, BigQuery and Kafka, and that fragmentation showed up twice over. When someone in the business needed a dataset, they'd ask around in person or over Slack, and wait anywhere from five minutes to an hour for the data team to point them to it. And any agent was left reasoning over whatever data it happened to be able to reach. It was a constant tax on the people waiting and the engineers pulled off their own work to answer, and a hard ceiling on how far Dojo could trust AI with real work.

So Dojo made DataHub the context layer over that estate. Now anyone can find a data asset by searching the catalog, or just by asking an AI agent in a Slack channel, connected through DataHub's MCP, see who owns it, and go straight to them. The same trusted, governed context the people use, the agents use too.

The results compound. The discovery tax largely disappears: business users find what they need in seconds instead of an hour, and engineers get their time back. Because people can finally see what already exists, the organization stops quietly creating duplicate copies of the same data, and the visibility cuts the other way, letting Dojo spot old, unowned systems nobody uses any more and retire them with confidence. The payoff is a team that moves faster, on data, and AI it can actually trust.

Know your data, trust your AI

Dojo is one example, other teams put DataHub to work in their own ways. A software company uses column-level lineage to answer "if I change this Kafka topic or table column, what breaks?" before they ship the change, not after. A marketing-tech company is building a single context layer across all of its systems. And a fast-scaling consumer company rolled DataHub out to everyone in the organization, every team and every agent, not just a licensed few, because with no per-seat tax, there was no reason not to.

Underneath all of them are the same two things a catalog is really for: knowing what data you have, and trusting what your AI does with it.

Know your data:

  • Discover. Search across every dataset, table, column, dashboard and pipeline in seconds, or let an agent do it for you. DataHub spans your whole estate, not just Aiven: connect Aiven services in a few clicks, and any other source like BigQuery, Snowflake, dbt and more through its ingestion framework.
  • Understand and trust. Column-level lineage shows where data comes from and what it feeds, with ownership and definitions attached. Classification and governance, including PII, mean you don't just find data, you know whether you can trust it and who's allowed to use it.

Trust your AI:

  • Agent-ready context. DataHub exposes the catalog to your agents through an MCP server. Before an agent picks a table or writes a query, it asks DataHub what data exists, what each field actually means, where it came from, who owns it and what governance rules apply, then grounds its next step in that metadata instead of guessing. That's the difference between an agent that quietly confuses revenue with profit and one that doesn't: trusted context in, trustworthy answers out

Two things set DataHub apart from a traditional catalog. It’s open and fully managed, genuine open source DataHub core, not a proprietary fork, so your metadata stays yours.

And it comes with unlimited users and no per-seat licensing, which is what actually makes a catalog work. Most catalog rollouts fail on a no-win choice: either you pay to license the whole company, including people who'll never open it, or you ration seats to control cost and adoption collapses. A catalog only delivers value when everyone can use it, so either way it underdelivers. Aiven removes the trade-off, everyone's in, from data engineers and analysts to agents and your CEO, so the catalog actually earns its keep.

What's new in GA

We spent Limited Availability running guided PoCs to get the experience right across real data estates. Here's what changes today:

  • Backed by an SLA. DataHub now runs with a 99.9% uptime SLA and Aiven support behind it.
  • Tested well beyond production load. We stress-tested DataHub well beyond production: around 100 agents and users searching and querying it at once, while large ingestion jobs ran at the same time. Entity lookups stayed under 100 ms with zero failures at every level tested, and ingestion kept running cleanly with zero errors (~100 assets per second, about a million assets in three hours). And on top of that, the underlying infrastructure scales on demand: when you need more headroom, add capacity without rebuilding.
  • Scales on demand if you need more headroom, add capacity without rebuilding.
  • Secure and maintained. Our team monitors for vulnerabilities and applies patches for you, so the catalog stays current and secure with no work on your side.
  • Seamless to operate. Version upgrades (DataHub is currently on v1.6), scaling, and even migration to another cloud are handled by Aiven and monitored around the clock; no maintenance windows to plan, no operators to hire.

Works across your whole estate, and better on Aiven

Aiven DataHub catalogs the data you already have, wherever it lives. With over 100 connectors, it ingests from BigQuery, Snowflake, dbt, MongoDB, your databases and much more, one place to find, understand and trust data across every system, no matter who runs it.
If your data is on Aiven, it goes further, and it's effortless. Connecting an Aiven service is as simple as selecting it; there's no connector to configure. As the context and governance layer of the Aiven Data & AI Cloud, DataHub automatically ingests what's there and builds lineage across your services - a single, live picture of your Kafka, PostgreSQL, ClickHouse, OpenSearch and MySQL data and how it flows between them, with no manual wiring.

Either way, the payoff is the same shape: DataHub gives your agents trusted context; Aiven Runtime gives them somewhere to run, next to the data; and MCP gives them a standard way to reach both. Fresh, governed data and somewhere to act on it, on open source, on any cloud.

Get started

Aiven DataHub is generally available today. You can have a catalog running, your data discoverable, your lineage mapped, your agents grounded in minutes, not months.

Request Access

Want to see it first? Watch the 2-minute walkthrough or book a demo.
Give your data and your agents the context they deserve.