TL;DR
If most of your Snowflake or Databricks spend goes to dashboards and reports, you are paying for a platform built for much bigger problems. Aiven for ClickHouse runs those workloads on a fixed plan, so adding dashboard users does not add to your compute bill. Native integrations with PostgreSQL and Apache Kafka also replace most of the ingestion and orchestration tools around your current platform. Move one dashboard at a time, and keep Snowflake or Databricks for work such as model training.
If most of your Snowflake or Databricks spend goes to dashboards and reports, you are paying for a platform built for much bigger problems. Snowflake and Databricks are good at what they were built for: large-scale transformation, machine learning and governance across hundreds of teams. Plenty of organizations adopted them for a much narrower job, which is getting data out of their application databases and event streams and into dashboards. For those organizations, the bill grows with every new dashboard and every new user, while the workload itself stays simple.
This post is for builders in that position. It covers how to tell whether your workload fits a simpler architecture, what that architecture looks like on the Aiven platform, and how to move over without a risky cutover.
Signs you are overpaying for dashboards
A few signals suggest your workload is simpler than the platform you run it on. Most of your compute goes to BI dashboards and scheduled reports rather than notebooks or model training. Your data is mostly structured and comes from PostgreSQL or MySQL tables and event streams. Your transformations are aggregations and joins that fit comfortably in SQL.
The cost pattern is the other signal. Your bill rises with the number of people looking at dashboards, and compute stays running through the working day because dashboards keep querying it. If most of this sounds familiar, a real-time analytical database is a better fit for the job.
How dashboards get expensive
Snowflake bills for a virtual warehouse while it runs, and Databricks follows a similar consumption model. That works well for batch jobs, where compute starts, finishes in minutes and suspends. Dashboards behave differently. They refresh on a schedule, many people open them at once, and each refresh wakes up a warehouse that then stays running. Supporting more concurrent users usually means a larger warehouse or more clusters, so cost tracks how many people look at the data more than how much work the queries do. Serverless options promise to remove the work of sizing and managing warehouses, but they still meter usage, so dashboards that query all day turn into spend that runs all day.
The rest of the stack adds to this. A typical setup pairs the platform with an ingestion tool to copy data out of the application database, an orchestrator to schedule transformations, and a transformation framework on top. Each of those is another contract, another integration and another system to operate. For a reporting workload, most of that stack exists to copy tables out of the application database and aggregate them for dashboards.
A database built for dashboard workloads
ClickHouse is an open source columnar database designed for this pattern: fast aggregations over large volumes of structured data, with many queries running at once. It stores data by column, compresses it heavily and reads only the columns a query needs, which keeps dashboard queries fast as data grows.
Aiven for ClickHouse, one of the fastest growing services on the Aiven platform, runs as a fully managed service on dedicated compute in the cloud and region of your choice. The plan you choose sets the service size, so more people opening a dashboard does not translate into more compute billing.
Connect your data without an ingestion tool
The bigger advantage comes from the platform around the database. On Aiven, the services that hold your source data connect to Aiven for ClickHouse through native integrations, so most of the ingestion layer in a typical Snowflake or Databricks stack goes away.
Build dashboard tables from PostgreSQL
Most organizations in this position start with an application database. Aiven for PostgreSQL connects to Aiven for ClickHouse as a managed integration. You enable it in the Aiven Console, and your PostgreSQL tables become queryable from ClickHouse without writing connection code or handling credentials. From there, a refreshable materialized view turns a PostgreSQL table into the rollup your dashboard reads and re-runs it on a schedule:
Loading code...
The example assumes a PostgreSQL service named app-pg with an orders table in the public schema of a database named sales, and an integration configured for that database. The integration uses defaultdb unless you choose another database when you set it up. Run the statements as the avnadmin user, since Aiven for ClickHouse only allows that user to create databases in SQL.
Every hour, ClickHouse re-runs the query against PostgreSQL and replaces the contents of the rollup, so the dashboard always reads a small, pre-aggregated table. Chaining several of these views with the DEPENDS ON clause covers a lot of what organizations use an orchestrator and a dbt model for today, and we covered that pattern in Replacing cron jobs and dbt pipelines with ClickHouse Refreshable Materialized Views.
Each refresh reads the full source table from PostgreSQL. That works well for small and medium tables, but as tables grow it adds load to your application database. For large or fast-changing tables, stream the changes through Apache Kafka instead.
Stream events and database changes through Kafka
For event data, and for change data capture from larger tables, Aiven for Apache Kafka connects to Aiven for ClickHouse as a managed integration. You choose the topics, the integration creates Kafka engine tables for them, and a materialized view writes each batch of messages into a MergeTree table as it arrives. To capture changes from PostgreSQL, MySQL or Microsoft SQL Server, run a Debezium source connector on Aiven for Apache Kafka Connect and point the integration at the resulting topics.
Load history from GCS, S3 and Azure Blob Storage
Other sources connect through managed credentials integrations. Aiven stores the connection details for GCS, Azure Blob Storage or Amazon S3 inside ClickHouse, so you can reference them by name in queries. This lets you load history without putting secrets in your SQL. Unload your historical tables from Snowflake or Databricks to GCS, Azure Blob Storage or S3 as Parquet, then read them into Aiven for ClickHouse in a single INSERT INTO ... SELECT.
Keep years of history on lower-cost storage
Dashboards often need years of history even though most queries only touch the last few weeks. Tiered storage in Aiven for ClickHouse keeps frequently queried data on fast block storage and moves data to object storage, either when block storage reaches 80% of its capacity or after a TTL you define on the table. Data in object storage stays queryable with the same SQL, and Aiven for ClickHouse caches it locally when it is read.
Keep your data catalog with DataHub
Moving off your current platform raises a fair question about governance. Aiven for DataHub connects to the Aiven services in your organization and builds lineage across them, so you can trace a number on a dashboard from Aiven for ClickHouse back through Apache Kafka to the PostgreSQL table it came from. Each service you add extends that map, and external systems such as the Snowflake or Databricks account you keep for heavier work can also join the same catalog.
Around all of this, the Aiven platform gives you one console, single sign-on and consolidated billing across every service, and each one is covered by the same compliance program, including ISO 27001, PCI DSS and an ISAE 3000 Type 2 report aligned with SOC 2.
Migrate one dashboard at a time
You do not need to move everything, and you do not need a big-bang migration. Keep Snowflake or Databricks for work such as model training or large Spark jobs, and move the high-concurrency dashboard traffic to Aiven for ClickHouse, since that traffic drives the cost pattern described above. Start with the dashboard that costs the most to run and connect its sources through the PostgreSQL or Kafka integration. If you already use dbt, point your existing models at Aiven for ClickHouse and adjust the parts that rely on platform-specific SQL. For simpler pipelines, refreshable materialized views handle the scheduling inside the database. Either way, have the dashboard read from pre-aggregated tables rather than raw tables, which keeps its queries fast and predictable as data grows. Most BI tools connect to ClickHouse directly, including Power BI, Tableau, Superset, Metabase and Looker, so you can point a copy of the dashboard at Aiven for ClickHouse and compare the results side by side.
Once the numbers match, switch the dashboard over and repeat with the next one. Each step moves more load off your current platform, and you can stop wherever the split makes sense for your organization.
Talk to an Aiven ClickHouse expert
If a large share of your Snowflake or Databricks spend goes to dashboards, we would be glad to walk through your workload with you. Book a session with an Aiven ClickHouse expert and we will map your current sources, transformations and dashboards onto the Aiven platform. If you would rather start on your own, deploy Aiven for ClickHouse and connect your first PostgreSQL service.
Table of contents
- Signs you are overpaying for dashboards
- How dashboards get expensive
- A database built for dashboard workloads
- Connect your data without an ingestion tool
- Build dashboard tables from PostgreSQL
- Stream events and database changes through Kafka
- Load history from GCS, S3 and Azure Blob Storage
- Keep years of history on lower-cost storage
- Keep your data catalog with DataHub
- Migrate one dashboard at a time
- Talk to an Aiven ClickHouse expert

