Oct 10, 2026

Observability is a data problem

Matching metrics, logs, and traces to the right backend on Aiven.

Tony Piazza |

RSS Feed

Senior Solutions Architect

TL;DR

Metrics, logs, and traces each have a different shape, so each needs a different backend. Aiven for Metrics handles metrics and alerting, Aiven for OpenSearch handles keyword search across logs, and Aiven for ClickHouse handles analysis of high-volume, high-cardinality events. All three are upstream open source engines on the same Aiven platform, so you can start with one and add the others as your needs change.

When I mention ClickHouse in a conversation about observability, prospects and customers are often surprised. Some people think of observability as a product category, something you buy from a vendor as one agent, one set of dashboards, and one bill. In that picture ClickHouse is an analytics database, so it seems out of place. The surprise goes away once you treat observability as what it is: a data problem. Your systems emit telemetry, which is just data, and each kind needs a different backend.

This post covers what observability means at a high level, the three kinds of telemetry most organizations collect, and the three Aiven services built for them: Aiven for Metrics, Aiven for OpenSearch®, and Aiven for ClickHouse®.

What observability means

Monitoring tells you whether the things you planned to watch are healthy. Observability is the ability to answer questions you did not plan for, using the data your systems already emit. When a checkout page slows down for some customers in one region after a deploy, nobody wrote a dashboard for that exact case ahead of time. You need data detailed enough, and a way to query it fast enough, to work it out.

That telemetry usually takes three shapes: metrics, logs, and traces. Metrics are numbers over time, such as request rate, error rate, and latency percentiles. Logs are records of individual events, often semi-structured text. Traces follow a single request across services, and in practice they are structured events with timing attached. Each shape is stored and queried differently, so no single backend is the best fit for all of them.

Three kinds of telemetry, three backends

Aiven for Metrics: metrics and alerting

Aiven for Metrics is built on Thanos, the Cloud Native Computing Foundation (CNCF) project that extends Prometheus with long-term storage and the ability to query many Prometheus servers through a single endpoint. It is compatible with Prometheus, so you can keep your existing PromQL queries, alerting rules, and Grafana® dashboards. Thanos keeps data in object storage and downsamples older data, so months or years of history stay affordable to store and fast to query at a coarser resolution.

This is the right home for the question "is something wrong right now, and how does it compare to last month?" It is less suited to high-cardinality data. A label that holds a device ID or request ID creates a new time series for every value, and that is the most common way a metrics store becomes slow and expensive.

Aiven for OpenSearch: searching logs

OpenSearch is a search engine, and once an alert fires, search is usually the next step in an investigation: find every log line containing this error, this host, or this customer, across every service, in the last hour. OpenSearch indexes text for exactly that, and OpenSearch Dashboards gives you a place to explore the results. Aiven services can also send their own logs to an Aiven for OpenSearch service through a built-in integration.

This is the right home for "what went wrong, and where?" It also fits security and audit logs, where finding the specific record matters more than aggregating millions of them. The trade-off is that full-text indexing costs storage and compute. At very high ingest volumes, or for heavy aggregations over long time ranges, the cost grows quickly; however, you can mitigate these costs by using hot-warm tiering to move older data to more economical storage.

Aiven for ClickHouse: analyzing high-volume events

This is where the surprise usually comes in. ClickHouse is a columnar analytics database, built for very large tables with a lot of columns. Once structured, that is exactly what logs and traces are. A columnar store reads only the columns a query touches and compresses each column well, because values within a column tend to repeat. As a result, ClickHouse stores large volumes of telemetry in a fraction of the space and aggregates billions of rows in seconds.

High cardinality is not a problem either. Customer ID, region, build version, and endpoint are ordinary columns, so you group by whichever ones the question needs, in SQL. Consider a question that is awkward in a metrics store and slow in a search engine: which ten customers saw the worst p99 latency this week?

Loading code...

Aiven for ClickHouse is the right home for "why is it happening, and for whom?" It is also where organizations tend to move log and trace data once volume has made cost the main topic of conversation. It is not a replacement for full-text search, and it is not a PromQL alerting stack, so it works best alongside the other two.

ClickHouse offers superior performance but requires trade-offs in developer overhead: it lacks native alerting and a specialized UI, shifting the responsibility for schema optimization and query design (via SQL) to the user.

Side by side, the three look like this:

SignalTypical questionBest fit
Metrics and alertsIs something wrong right now?Aiven for Metrics
Logs you search by textWhat went wrong, and where?Aiven for OpenSearch
High-volume, high-cardinality eventsWhy is it happening, and for whom?Aiven for ClickHouse

How the pieces fit together

Most organizations end up using two or three of these, so the real design decision is routing. A common pattern is to collect telemetry with OpenTelemetry, publish it to Apache Kafka as a buffer, and send each signal to the backend that fits it. Kafka absorbs ingest spikes during an incident, which is exactly when telemetry volume jumps, and it lets you add or change a destination without touching the applications that produce the data. Grafana sits on top as a single place to view metrics from Thanos, logs from OpenSearch, and events from ClickHouse.

Aiven unifies the entire pipeline into a single platform with consolidated billing and management. Built-in connectors stream Kafka data directly into ClickHouse and OpenSearch, while Aiven for Grafana provides a centralized visualization layer across all three backends.

Why upstream open source matters

Telemetry is long-lived and high-volume, which makes it one of the most expensive data sets to move once it is in place. That is why it matters which engine actually runs underneath your observability stack.

Many vendors that build an open source project run something different in their own cloud, such as a proprietary storage engine or features that never ship in the open source release. The queries look the same at first, but over time your schemas, settings, and operational habits come to depend on behavior you cannot run anywhere else. Aiven runs the upstream open-source engines, ensuring your architecture is built on standard ClickHouse, OpenSearch, and Thanos.

The projects behind this stack reflect the same concern. OpenSearch exists because Elasticsearch moved away from an open source license, and Thanos is governed by the CNCF rather than by a single company.

Running upstream engines also makes local testing simple. You can prototype on a laptop with the same ClickHouse, OpenSearch, or Thanos you’ll run in production, using the official open source Docker images. When you’re ready, point your queries and dashboards at Aiven, usually with no changes.

Building observability on the Aiven platform

If you already run Prometheus, start with Aiven for Metrics and keep your dashboards and alerts. If you mostly search your logs by keyword, add Aiven for OpenSearch. When log or trace volume, or the number of dimensions you need to slice by, starts to drive up cost and slow down queries, move that data into Aiven for ClickHouse.

You do not have to pick one backend for everything. On Aiven, you can start with a single service and add the others on the same platform as your telemetry grows. As an added benefit, you can use Aiven for Grafana to query all of these from the same dashboards.

If you are deciding where your own telemetry should live, we would be glad to talk through the options with you.